Paper Overview
- Field: NLP (mathematical reasoning for LLMs)
- Authors: Aozhe Wang, Zhengxi Lu, Jianze Wang
- arXiv: 2508.11369
- Asymmetric objective: distill rollouts consistent with the majority-vote pseudo-label; apply Grouped RL penalties only to inconsistent rollouts, which are typically wrong even when the vote is wrong.
- Token-level selection: distillation down-weights already-converged positions, while RL penalizes only confident errors—keeping both updates well-defined despite frequent pseudo-label mistakes.
- Tighter self-supervision: majority-vote routing yields progressively tighter supervision as the model improves.
- Results:
- Matches label-supervised OPSD on five competition-level benchmarks without any labels.
- Improves Qwen3-1.7B from 38.0% to 45.2% in test-time training.
- Gains of +25.2% to +36.4% in non-thinking mode.
- Strong cross-task generalization.
Abstract (English)
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. The authors observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct.
Building on this observation, they propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL.