OALib Journal期刊
ISSN: 2333-9721
费用：99美元

投递稿件

查看量	下载量

相关文章
更多...

- 2018

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation
Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

DOI: 10.15918/j.jbit1004-0579.201827.0116

Huu-anh Tran,Yuhang Guo,Ping Jian,Shumin Shi,Heyan Huang

Keywords: parallel corpus filtering low resource languages bilingual movie subtitles machine translation Chinese-Vietnamese translation
parallel corpus filtering low resource languages bilingual movie subtitles machine translation Chinese-Vietnamese translation

Full-Text Cite this paper Add to My Lib

Abstract:

The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource.However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed:sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.
The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource.However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed:sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.

Full-Text

Contact Us

service@oalib.com

QQ:3279437679

WhatsApp +8615387084133

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine TranslationImproving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation
Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation