دانلود رایگان مقاله: جستجو برای کلمات تفاوتی در فضای مداوم چند بعدی

کد مقاله	کد نشریه	سال انتشار	مقاله انگلیسی	نسخه تمام متن
11002933	1451674	2019	45 صفحه PDF	دانلود رایگان

عنوان انگلیسی مقاله ISI

Searching for discriminative words in multidimensional continuous feature space

ترجمه فارسی عنوان

جستجو برای کلمات تفاوتی در فضای مداوم چند بعدی

دانلود مقاله + سفارش ترجمه

دانلود مقاله ISI انگلیسی

رایگان برای ایرانیان

کلمات کلیدی

Text categorisation NLP Feature vectors - بردارهای ویژگی Distributed representation - نمایندگی توزیع شده

موضوعات مرتبط

مهندسی و علوم پایه مهندسی کامپیوتر پردازش سیگنال

پیش نمایش مقاله

جستجو برای کلمات تفاوتی در فضای مداوم چند بعدی

چکیده انگلیسی

Word feature vectors have been proven to improve many natural language processing tasks. With recent advances in unsupervised learning of these feature vectors, it became possible to train it with much more data, which also resulted in better quality of learned features. Since it learns joint probability of latent features of words, it has the advantage that we can train it without any prior knowledge about the goal task we want to solve. We aim to evaluate the universal applicability property of feature vectors, which has been already proven to hold for many standard NLP tasks like part-of-speech tagging or syntactic parsing. In our case, we want to understand the topical focus of text documents and design an efficient representation suitable for discriminating different topics. The discriminativeness can be evaluated adequately on text categorisation task. We propose a novel method to extract discriminative keywords from documents. We utilise word feature vectors to understand the relations between words better and also understand the latent topics which are discussed in the text and not mentioned directly but inferred logically. We also present a simple way to calculate document feature vectors out of extracted discriminative words. We evaluate our method on the four most popular datasets for text categorisation. We show how different discriminative metrics influence the overall results. We demonstrate the effectiveness of our approach by achieving state-of-the-art results on text categorisation task using just a small number of extracted keywords. We prove that word feature vectors can substantially improve the topical inference of documents' meaning. We conclude that distributed representation of words can be used to build higher levels of abstraction as we demonstrate and build feature vectors of documents. Our method can help in any multi-domain environment to automatically extract discriminative keywords. It can be used to organise and search documents more efficiently.

ناشر

Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Computer Speech & Language - Volume 53, January 2019, Pages 276-301

نویسندگان

Marius Sajgalik, Michal Barla, Maria Bielikova,

علوم انسانی و هنر

فنی، مهندسی و علوم پایه

پزشکی و سلامت

بیو تکنولوژی

پذیرش سفارش ترجمه

دانلود رایگان مقاله ISI : جستجو برای کلمات تفاوتی در فضای مداوم چند بعدی

دسترسی سریع

ارتباط

English Website