Large-Coverage Root Lexicon Extraction for Hindi

Cohan Sujay Carlos, Monojit Choudhury, and Sandipan Dandapat

Abstract

This paper describes a method using morphological rules and heuristics, for the automatic

extraction of large-coverage lexicons of stems and root word-forms from a raw text corpus. We cast the problem of high-coverage lexicon extraction as one of stemming followed by root word-form selection. We examine the use of POS tagging to improve precision and recall of stemming and thereby the coverage of the lexicon. We present accuracy, precision and recall scores for the system on a Hindi corpus.

Details

Publication typeInproceedings
Published inProceedings of EACL 2009
URLhttp://www.aclweb.org/anthology/E/E09/E09-1015.pdf
Pages121 - 129
PublisherAssociation for Computational Linguistics
> Publications > Large-Coverage Root Lexicon Extraction for Hindi