کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
5103372 1480104 2017 20 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Multifractal correlations in natural language written texts: Effects of language family and long word statistics
ترجمه فارسی عنوان
همبستگی مولتی فکتال در متون نوشته شده به زبان طبیعی: تأثیر خانواده زبان و آمار کلام طولانی
کلمات کلیدی
چند فاکتوریل، شمارش جعبه، همبستگی های طولانی مدت، زبان، خانواده های زبان، متون نوشته شده،
موضوعات مرتبط
مهندسی و علوم پایه ریاضیات فیزیک ریاضی
چکیده انگلیسی
During the last years, several methods from the statistical physics of complex systems have been applied to the study of natural language written texts. They have mostly been focused on the detection of long-range correlations, multifractal analysis and the statistics of the content word positions. In the present paper, we show that these statistical aspects of language series are not independent but may exhibit strong interrelations. This is done by means of a two-step investigation. First, we calculate the multifractal spectra using the word-length representation of huge parallel corpora from ten European languages and compare with the shuffled data to assess the contribution of long-range correlations to multifractality. In the second step, the detected multifractal correlations are shown to be related to the scale-dependent clustering of the long, highly informative content words. Furthermore, exploiting the language sensitivity of the used word-length representation, we demonstrate the consistent impact of the classification of languages into families on the multifractal correlations and long-word clustering patterns.
ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Physica A: Statistical Mechanics and its Applications - Volume 469, 1 March 2017, Pages 173-182
نویسندگان
, , , , , , ,