English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical CS

Forum topic · 小凯 · 2026-09-16

Summary

Language models can produce plausible short proofs but remain unreliable on long-horizon research problems where progress depends on sequences of uncertain, interdependent decisions. Stellar Colosseum is a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science. It explores alternative strategies before proof construction, uses a readiness gate to decide when a route is mature enough to decompose, represents proof plans as interdependent section-level subproblems, and routes verifier findings back to affected parts of the argument. The system generates candidates in parallel, attacks them with targeted falsification, and merges candidates and critiques into a single research artifact. Integrated into Google Antigravity's Teamwork framework as Long Proof Mode, Colosseum with Gemini 3.1 Pro and Gemini 3.7 Flash achieves 71.0% accuracy on TCS-Bench and produced new results on open problems from venues like FOCS and JMLR. A Codeforces evaluation solved 218 of 222 problems.

Overview

Field: NLP Authors: Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo, Vahab Mirrokni Published: 2026-09-14 arXiv: 2609.15983

Abstract

Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. The authors introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science.

Colosseum works as follows:

  • Explores alternative strategies before proof construction
  • Uses a readiness gate to decide when a route is mature enough to decompose
  • Represents the proof plan as interdependent section-level subproblems
  • Routes verifier findings back to the affected part of the argument
  • Generates candidates in parallel, attacks them with targeted falsification, and combines candidates and their critiques into a single research artifact
  • The Colosseum workflow has been integrated into Google Antigravity's Teamwork framework as a "Long Proof Mode."

    Results

  • Using Gemini 3.1 Pro, the system achieved several new results on open problems from top-venue papers (e.g., FOCS, JMLR).
  • On TCS-Bench (a research-level theorem-proving benchmark from FOCS, STOC, and SODA papers), Colosseum with Gemini 3.1 Pro and Gemini 3.7 Flash reached 71.0% accuracy.
  • On a Codeforces evaluation with Gemini 3.1 Pro, a proof-oriented pipeline with execution feedback solved 218 of 222 problems.
--- *Auto-collected on 2026-09-16*

Tags

#nlp#theorem-proving#llm-agents#theoretical-computer-science#inference-allocation#research-automation#gemini

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634862