Article ID Journal Published Year Pages File Type
1110926 Procedia - Social and Behavioral Sciences 2015 5 Pages PDF
Abstract

Over the last few decades, a large evolution has been made in the field of handwritten recognition. Material of handwritten documents is become less with current trends of digital electronics. However, for the investigation and research on a particular language a large volume of handwritten documents database is required. In this paper we describe our approach for development a large volume of Urdu handwritten text images Corpus on Urdu language. To make the database available in large field of Natural Language Processing we annotate database for each image and associate a XML based ground-truth Meta information to make it computer compatible as a linguistic resource. This paper focus on the some issue related with Corpus design and annotation such as data collection, writers selection, methodology of annotation etc.

Related Topics
Social Sciences and Humanities Arts and Humanities Arts and Humanities (General)