کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
490242 705691 2014 10 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
A Workflow Application for Parallel Processing of Big Data from an Internet Portal
ترجمه فارسی عنوان
یک برنامه گردش کار برای پردازش موازی داده های بزرگ از یک پورتال اینترنتی یک ؟؟ یک ؟؟
موضوعات مرتبط
مهندسی و علوم پایه مهندسی کامپیوتر علوم کامپیوتر (عمومی)
چکیده انگلیسی

The paper presents a workflow application for efficient parallel processing of data downloaded from an Internet portal. The workflow partitions input files into subdirectories which are further split for parallel processing by services installed on distinct computer nodes. This way, analysis of the first ready sub-directories can start fast and is handled by services implemented as parallel multithreaded applications using multiple cores of modern CPUs. The goal is to assess achievable speed-ups and determine which factors influence scalability and to what degree. Data processing services were implemented for assessment of context (positive or negative) in which the given keyword appears in a document. The testbed application used these services to determine how a particular brand was recognized by either authors of articles or readers in comments in a specific Internet portal focused on new technologies. Obtained execution times as well as speed-ups are presented for data sets of various sizes along with discussion on how factors such as load imbalance and memory/disk bottlenecks limit performance.

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Procedia Computer Science - Volume 29, 2014, Pages 499-508