An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
研究领域: ML 作者: Ivan Moshkov, Stephen Ge, George Armstrong 发布时间: 2026-09-11 arXiv: 2509.05823
论文概要
研究领域: ML 作者: Ivan Moshkov, Stephen Ge, George Armstrong 发布时间: 2026-09-11 arXiv: 2509.05823
中文摘要
我们研究模型后训练和测试时推理设计如何影响困难奥赛数学的自然语言证明生成。从Nematron 3 Ultra出发,我们使用监督微调和强化学习训练两个专家检查点,并评估检查点选择、验证和优化。基于这些发现,我们提出开放模型测试时计算管线。系统完全以自然语言运行,无形式化证明器、外部工具或互联网访问。三个Nematron 3 Ultra检查点——通用可用模型和两个后训练专家——驱动迭代搜索,生成、验证和优化候选证明;独立的强计算阶段随后选择每份最终提交。系统在IMO 2026上获得42分中的30分,达到金牌门槛。我们发布两个后训练检查点以及训练数据、训练和推理代码、提交的解决方案,还有Nematron-IMO-Bench——包含200道新奥赛级别问题的新基准。
原文摘要
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points a...
*自动采集于 2026-09-12*
#论文 #arXiv #ML #小凯