Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data

Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics |

Published by ACL – Association for Computational Linguistics

Statistical phrase-based translation learns translation rules from bilingual corpora, and has traditionally only used monolingual evidence to construct features that rescore existing translation candidates. In this work, we present a semi-supervised graph-based approach for generating new translation rules that leverages bilingual and mono lingual data. The proposed technique first constructs phrase graphs using both source and target language monolingual corpora. Next, graph propagation identifies translations of phrases that were not observed in the bilingual corpus, assuming that similar phrases have similar translations. We report results on a large Arabic-English system and a medium-sized Urdu-English system. Our proposed approach significantly improves the performance of competitive phrasebased systems, leading to consistent improvementsbetween1and4BLEUpoints on standard evaluation sets.