کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
493016 721666 2013 8 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Arabic Character Recognition System Development
موضوعات مرتبط
مهندسی و علوم پایه مهندسی کامپیوتر علوم کامپیوتر (عمومی)
پیش نمایش صفحه اول مقاله
Arabic Character Recognition System Development
چکیده انگلیسی

We develop Arabic Optical Character Recognition (AOCR) system that has five stages: preprocessing, segmentation, thinning, feature extraction, and classification. In preprocessing stage, we compare two skew estimation algorithms i.e. skew estimation by image moment and by skew triangle. We also implemented binarization and median filter. In thinning stage, we use Hilditch thinning algorithm incorporated by two templates, one to prevent superfluous tail and the other one to remove unnecessary interest point. In segmentation stage, line segmentation is done by horizontal projection cross verification by standard deviation, sub-word segmentation is done by connected pixel components, and letter segmentation is done by Zidouri algorithm. In the feature extraction stage, 24 features are extracted. The features can be grouped into three groups: main body features, perimeter- skeleton features, and secondary object features. In the classification stage, we use decision tree that generated by C4.5 algorithm. Functionality test showed that skew estimation using moment is more accurate than using skew triangle, median filter tends to erode the letter shape, and template addition into Hilditch algorithm gives a good result. Performance test yield these result. Line segmentation had 99.9% accuracy. Standard deviation is shown can reduce over-segmentation and quasi-line. Letter segmentation had 74% accuracy, tested on six different fonts. Classification components had 82% accuracy, tested by cross validation. Unfortunately, overall performance of the system only reached 48.3%.

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Procedia Technology - Volume 11, 2013, Pages 334-341