A Structure for Annotation and Ground-truthing of Urdu Handwritten Text Image Corpus

Article ID	Journal	Published Year	Pages	File Type
1110926	Procedia - Social and Behavioral Sciences	2015	5 Pages	PDF

Abstract

Over the last few decades, a large evolution has been made in the field of handwritten recognition. Material of handwritten documents is become less with current trends of digital electronics. However, for the investigation and research on a particular language a large volume of handwritten documents database is required. In this paper we describe our approach for development a large volume of Urdu handwritten text images Corpus on Urdu language. To make the database available in large field of Natural Language Processing we annotate database for each image and associate a XML based ground-truth Meta information to make it computer compatible as a linguistic resource. This paper focus on the some issue related with Corpus design and annotation such as data collection, writers selection, methodology of annotation etc.