A Kumaran, Ranbeer Makin, Vijay Pattisapu, Shaik Sharif, Gary Kacmarcik, and Lucy Vanderwende
Automatic extraction of semantic information, if successful, offers to
languages with little or poor resources, the prospects of creating ontological resources inexpensively, thus providing support for commonsense reasoning applications in those languages. In this paper we explore the automatic extraction of synonymy information from large corpora using two complementary techniques: a generic broad-coverage parser for generation of bits of semantic information, and their synthesis into sets of synonyms using sense-disambiguation. To validate the quality of the synonymy information thus extracted, we experiment with English, where appropriate semantic resources are already available. We cull synonymy information from a large corpus and compare it against synonymy information available in several standard sources. We present the results of our methodology, both quantitatively and qualitatively; we show that good quality synonymy information may be extracted automatically from large corpora using the proposed methodology.
|Published in||the Ontologies in Text Technology Workshop, Osnabruck, Germany|