Paper Overview
Field: NLP Authors: Zhiqiu Xu, Shibo Jin, Shreya Arya arXiv: 2604.21937
Introduction
As frontier language models reach near-ceiling performance on static mathematical benchmarks, existing evaluations increasingly fail to differentiate model capabilities — largely because they treat models solely as solvers of fixed problem sets. MathDuels addresses this by introducing a self-play benchmark in which models occupy dual roles.
How It Works
- Dual roles: Each model authors math problems under adversarial prompting and also solves problems authored by every other participant.
- Three-stage generation pipeline: meta-prompting, problem generation, and difficulty amplification.
- Independent verifier: Validates generated problems and excludes ill-posed questions.
- Rasch model: Jointly estimates solver abilities and problem difficulties; author quality is derived from the difficulty of each model's authored problems.
- Evaluation across 19 frontier models shows that problem-posing and problem-solving abilities are only partially decoupled.
- Dual-role evaluation reveals capability differences that are invisible in single-role benchmarks.
- As new models enter the arena, their authored problems defeat previously dominant solvers — so benchmark difficulty co-evolves with participant strength rather than saturating at a fixed ceiling.
- arXiv: 2604.21937
Key Findings
Leaderboard
A public leaderboard is hosted and updated as new model releases arrive.
Links
*Auto-collected on 2026-04-25*