Statistical Significance Tests for Machine Translation Evaluation

Philipp Koehn

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract / Description of output

If two translation systems differ differ in performance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling methods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real.
Original languageEnglish
Title of host publicationProceedings of EMNLP 2004
EditorsDekang Lin, Dekai Wu
Place of PublicationBarcelona, Spain
PublisherAssociation for Computational Linguistics
Number of pages8
Publication statusPublished - 1 Jul 2004


Dive into the research topics of 'Statistical Significance Tests for Machine Translation Evaluation'. Together they form a unique fingerprint.

Cite this