کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
10355214 867112 2005 11 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Term norm distribution and its effects on Latent Semantic Indexing
موضوعات مرتبط
مهندسی و علوم پایه مهندسی کامپیوتر نرم افزارهای علوم کامپیوتر
پیش نمایش صفحه اول مقاله
Term norm distribution and its effects on Latent Semantic Indexing
چکیده انگلیسی
Latent Semantic Indexing (LSI) uses the singular value decomposition to reduce noisy dimensions and improve the performance of text retrieval systems. Preliminary results have shown modest improvements in retrieval accuracy and recall, but these have mainly explored small collections. In this paper we investigate text retrieval on a larger document collection (TREC) and focus on distribution of word norm (magnitude). Our results indicate the inadequacy of word representations in LSI space on large collections. We emphasize the query expansion interpretation of LSI and propose an LSI term normalization that achieves better performance on larger collections.
ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Information Processing & Management - Volume 41, Issue 4, July 2005, Pages 777-787
نویسندگان
, , ,