کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
859321 1470757 2014 8 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
A Fast Distributed Focused-web Crawling
ترجمه فارسی عنوان
سریع توزیع متمرکز-وب خزیدن یک ؟؟
موضوعات مرتبط
مهندسی و علوم پایه سایر رشته های مهندسی مهندسی (عمومی)
چکیده انگلیسی

Mining data from a web database becomes more challenging in recent years due to the exploding size of data, the rising of dynamic web, and the increasing performance of web security. Mining data from a web database differs from mining data from web sites because it is intended to collect specific data from a single web site. Collecting a very large data in a limited time tends to be detected as a cyber attack and will be banned from connecting into the web server. To avoid the problem, this paper proposes a crawling method to mine web database faster and cheaper than conventional web crawlers. The method used is to run hundreds of threads from a single web crawler in a single computer and to distribute the threads into hundreds or thousands publicly available proxy servers. This web crawler strategy highly increases the speed of mining and is more secure than using single thread of web crawler.

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Procedia Engineering - Volume 69, 2014, Pages 492-499