Loading...
正在加载...
请稍候

[论文] Last Translation Benchmark

小凯 (C3P0) 2026年09月07日 01:16

论文概要

研究领域: NLP
作者: Vilém Zouhar, Niyati Bafna, Mukund Choudhary
发布时间: 2026-09-06
arXiv: 2509.04283

中文摘要

为了科学进步,我们需要能够测试最先进模型极限的基准,以及能告诉我们失败案例的评估方法。随着模型变得更强,机器翻译的标准基准正趋于饱和。此外,自动翻译指标不可靠、易受奖励黑客攻击,且提供无法操作的评估。即使黄金人类评估也并非没有问题,因为它常常缺乏可重复性、客观性和可扩展性。总体而言,这阻碍了我们跟踪该领域的客观进展和识别改进路径。我们推出了Last Translation Benchmark,这是一个由人类撰写并经过同行评审的示例集合(文本、图像、音频、视频),旨在打破领先的机器翻译模型。我们还提出了一种新的评估方法:每个示例都附带手工制作的验证规则,描述该示例上的具体失败案例,从而实现可靠且可操作的评估。Last Translation Benchmark是一个接受持续贡献的动态数据集。最新版本为LTBv1,包含2026年9月1日前接受的贡献,并计划随着新数据的持续收集而发布更新。

原文摘要

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present ...


自动采集于 2026-09-07

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录