跳到主要导航 跳到搜索 跳到主要内容

Improved graph-based bilingual corpus selection with sentence pair ranking for statistical machine translation

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

In statistical machine translation, the number of sentence pairs in the bilingual corpus is very important to the quality of translation. However, when the quantity reaches some extent, enlarging corpus has less effect on the translation; whereas increasing greatly the time and space complexity to building translation systems, which hinders the development of statistical machine translation. In this paper, we propose several ranking approaches to measure the quantity of information of each sentence pair, and apply them into a graph-based bilingual corpus selection framework to form an improved corpus selection approach, which now considers the difference of the initial quantities of information between the sentence pairs. Our experiments in a Chinese-English translation task show that, selecting only 50% of the whole corpus via the graph-based selection approach as training set, we can obtain the near translation result with the one using the whole corpus, and we obtain better results than the baselines after using the IDF-related ranking approach.

源语言英语
主期刊名Proceedings - 2011 23rd IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2011
446-451
页数6
DOI
出版状态已出版 - 2011
活动23rd IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2011 - Boca Raton, FL, 美国
期限: 7 11月 20119 11月 2011

出版系列

姓名Proceedings - International Conference on Tools with Artificial Intelligence, ICTAI
ISSN(印刷版)1082-3409

会议

会议23rd IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2011
国家/地区美国
Boca Raton, FL
时期7/11/119/11/11

指纹

探究 'Improved graph-based bilingual corpus selection with sentence pair ranking for statistical machine translation' 的科研主题。它们共同构成独一无二的指纹。

引用此