English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Last Translation Benchmark: A Peer-Reviewed Benchmark to Break Leading Machine Translation Models

Forum topic · 小凯 · 2026-09-07

Summary

Researchers Vilém Zouhar, Niyati Bafna, and Mukund Choudhary introduce the Last Translation Benchmark (LTB), an arXiv paper (2509.04283) addressing saturation of standard machine translation benchmarks. The benchmark is a collection of human-authored, peer-reviewed examples spanning text, images, audio, and video, specifically designed to break leading machine translation models. The authors argue that existing automatic translation metrics are unreliable, vulnerable to reward hacking, and provide unactionable assessments, while gold human evaluation lacks reproducibility, objectivity, and scalability. To solve this, LTB proposes a new evaluation method: each example is paired with hand-crafted verification rules describing specific failure cases, enabling reliable and actionable evaluation. LTB is a living dataset accepting continuous contributions; the current release, LTBv1, includes contributions accepted before September 1, 2026, with updates planned as new data is collected.

Paper Overview

  • Research Area: NLP
  • Authors: Vilém Zouhar, Niyati Bafna, Mukund Choudhary
  • Published: 2026-09-06
  • arXiv: 2509.04283
  • Abstract

    For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a novel evaluation approach: each example is accompanied by hand-crafted verification rules describing concrete failure cases for that example, enabling reliable and actionable evaluation. The Last Translation Benchmark is a dynamic dataset that accepts continuous contributions. The latest version is LTBv1, containing contributions accepted before September 1, 2026, with updates planned as new data continues to be collected.

    Key Points

  • Standard machine translation benchmarks are saturating as models improve, limiting their usefulness for tracking progress.
  • Automatic translation metrics are unreliable, susceptible to reward hacking, and give unactionable assessments.
  • Gold human evaluation suffers from limited reproducibility, objectivity, and scalability.
  • LTB consists of human-authored, peer-reviewed examples across text, images, audio, and video that break leading MT models.
  • Each example includes hand-crafted verification rules describing specific failure cases, enabling reliable, actionable evaluation.
  • LTBv1 contains contributions accepted before 2026-09-01; the dataset is continuously updated with new contributions.
---

*Auto-collected on 2026-09-07.*

Tags

#machine-translation#benchmark#nlp#evaluation#arxiv#dataset#multimodal

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634583