Paper Overview
- Research Area: NLP
- Authors: Vilém Zouhar, Niyati Bafna, Mukund Choudhary
- Published: 2026-09-06
- arXiv: 2509.04283
- Standard machine translation benchmarks are saturating as models improve, limiting their usefulness for tracking progress.
- Automatic translation metrics are unreliable, susceptible to reward hacking, and give unactionable assessments.
- Gold human evaluation suffers from limited reproducibility, objectivity, and scalability.
- LTB consists of human-authored, peer-reviewed examples across text, images, audio, and video that break leading MT models.
- Each example includes hand-crafted verification rules describing specific failure cases, enabling reliable, actionable evaluation.
- LTBv1 contains contributions accepted before 2026-09-01; the dataset is continuously updated with new contributions.
Abstract
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a novel evaluation approach: each example is accompanied by hand-crafted verification rules describing concrete failure cases for that example, enabling reliable and actionable evaluation. The Last Translation Benchmark is a dynamic dataset that accepts continuous contributions. The latest version is LTBv1, containing contributions accepted before September 1, 2026, with updates planned as new data continues to be collected.
Key Points
*Auto-collected on 2026-09-07.*