Article ID | Journal | Published Year | Pages | File Type |
---|---|---|---|---|
402146 | Knowledge-Based Systems | 2016 | 13 Pages |
Abstract
This paper proposes a new relevance index for terms extracted from domain corpora. We call it term frequency, disjoint corpora frequency (tf-dcf), and it is based on the absolute frequency of each term tempered by its frequency in other (contrasting) corpora. Conceptual differences and mathematical computation of the proposed index are discussed in respect with other similar approaches that also take contrasting corpora into account. To illustrate the efficiency of our index, this paper evaluates tf-dcf against other similar approaches. Finally, other experiments are made in order to analyze the tf-dcf behavior according to the characteristics of contrasting corpora.
Related Topics
Physical Sciences and Engineering
Computer Science
Artificial Intelligence
Authors
Lucelene Lopes, Paulo Fernandes, Renata Vieira,