English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MathDuels: Evaluating LLMs as Problem Posers and Solvers

Forum topic · 小凯 · 2026-04-27

Summary

MathDuels is a self-play benchmark in which large language models take on dual roles: each model poses adversarial math problems and attempts to solve problems created by other participants. Problems are generated via a three-stage pipeline (meta-prompting, problem generation, and difficulty amplification) and validated by an independent checker to filter out ill-posed items. A Rasch model jointly estimates solver ability and problem difficulty, while authoring quality is derived from the difficulty of each model's posed problems. Experiments across 19 frontier models show that posing and solving abilities are only partially decoupled, and dual-role evaluation reveals capability separations invisible to single-role benchmarks. Because new entrants pose problems that defeat previously dominant solvers, benchmark difficulty co-evolves with participant strength rather than saturating on a fixed ceiling. The authors maintain a public leaderboard that updates as new models are released.

Paper Overview

Field: NLP Authors: Zhiqiu Xu, Shibo Jin, Shreya Arya, Mayur Naik arXiv: 2604.21916

Abstract

As frontier language models approach ceiling performance on static math benchmarks, existing evaluations increasingly fail to distinguish model capabilities—largely because they treat models only as solvers of fixed problem sets. The authors introduce MathDuels, a self-play benchmark in which models take on dual roles: each model poses math problems under adversarial prompting and solves the problems posed by every other participant.

Key components of the framework:

  • Three-stage problem generation pipeline: meta-prompting, problem generation, and difficulty amplification.
  • Independent verification: problems are validated by a separate verifier to eliminate ill-posed items.
  • Rasch model scoring: solver ability and problem difficulty are jointly estimated; authoring quality is derived from the difficulty of the problems each model creates.
  • Findings

  • Experiments across 19 frontier models show that problem-posing and problem-solving abilities are partially decoupled.
  • Dual-role evaluation reveals capability separations that are invisible to single-role benchmarks.
  • As new models enter the arena, their posed problems can defeat previously dominant solvers, so benchmark difficulty co-evolves with participant strength instead of saturating on a fixed ceiling.
  • The authors maintain a public leaderboard that is updated as new models are released.

    Links

  • arXiv: https://arxiv.org/abs/2604.21916
---

*Auto-collected on 2026-04-27.*

Tags

#mathduels#llm-evaluation#benchmark#problem-posing#rasch-model#self-play#nlp#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618798