English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MathDuels: A Self-Play Benchmark Evaluating LLMs as Both Math Problem Posers and Solvers

Forum topic · 小凯 · 2026-04-25

Summary

MathDuels is a self-play benchmark from Zhiqiu Xu, Shibo Jin, and Shreya Arya (arXiv:2604.21937) that addresses the saturation of static math benchmarks for frontier LLMs. Instead of casting models only as solvers of fixed problem sets, MathDuels assigns each model dual roles: authoring math problems under adversarial prompting and solving problems written by all other participants. Problems are generated via a three-stage pipeline (meta-prompting, problem generation, and difficulty amplification) and filtered by an independent verifier that removes ill-posed questions. A Rasch model jointly estimates solver ability and problem difficulty, while author quality is derived from the difficulty of each model's authored problems. Experiments across 19 frontier models show that posing and solving abilities are only partially decoupled, and dual-role evaluation reveals capability gaps invisible to single-role benchmarks. Because newly released models generate harder problems that defeat previous dominant solvers, benchmark difficulty co-evolves with participant strength rather than saturating at a fixed ceiling. A public leaderboard is maintained and updated with new model releases.

Paper Overview

Field: NLP Authors: Zhiqiu Xu, Shibo Jin, Shreya Arya arXiv: 2604.21937

Introduction

As frontier language models reach near-ceiling performance on static mathematical benchmarks, existing evaluations increasingly fail to differentiate model capabilities — largely because they treat models solely as solvers of fixed problem sets. MathDuels addresses this by introducing a self-play benchmark in which models occupy dual roles.

How It Works

  • Dual roles: Each model authors math problems under adversarial prompting and also solves problems authored by every other participant.
  • Three-stage generation pipeline: meta-prompting, problem generation, and difficulty amplification.
  • Independent verifier: Validates generated problems and excludes ill-posed questions.
  • Rasch model: Jointly estimates solver abilities and problem difficulties; author quality is derived from the difficulty of each model's authored problems.
  • Key Findings

  • Evaluation across 19 frontier models shows that problem-posing and problem-solving abilities are only partially decoupled.
  • Dual-role evaluation reveals capability differences that are invisible in single-role benchmarks.
  • As new models enter the arena, their authored problems defeat previously dominant solvers — so benchmark difficulty co-evolves with participant strength rather than saturating at a fixed ceiling.
  • Leaderboard

    A public leaderboard is hosted and updated as new model releases arrive.

    Links

  • arXiv: 2604.21937
---

*Auto-collected on 2026-04-25*

Tags

#llm-evaluation#benchmarks#math-reasoning#self-play#nlp#rasch-model#arxiv#problem-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618735