کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
8893813 1629383 2019 6 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Natural language indexing for pedoinformatics
موضوعات مرتبط
مهندسی و علوم پایه علوم زمین و سیارات فرآیندهای سطح زمین
پیش نمایش صفحه اول مقاله
Natural language indexing for pedoinformatics
چکیده انگلیسی
The multiple schema for the classification of soils rely on differing criteria but the major soil science systems, including the United States Department of Agriculture (USDA) and the international harmonized World Reference Base for Soil Resources soil classification systems, are primarily based on inferred pedogenesis. Largely these classifications are compiled from individual observations of soil characteristics within soil profiles, and the vast majority of this pedologic information is contained in non-quantitative text descriptions. We present initial text mining analyses of parsed text in the digitally available USDA soil taxonomy documentation and the Soil Survey Geographic database. Previous research has shown that latent information structure can be extracted from scientific literature using Natural Language Processing techniques, and we show that this latent information can be used to expedite query performance by using syntactic elements and part-of-speech tags as indices. Technical vocabulary often poses a text mining challenge due to the rarity of its diction in the broader context. We introduce an extension to the common English vocabulary that allows for nearly-complete indexing of USDA Soil Series Descriptions.
ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Geoderma - Volume 334, 15 January 2019, Pages 49-54
نویسندگان
, , ,