Paper Overview
- Field: NLP
- Authors: Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen
- Published: 2026-08-27
- arXiv: 2608.27448
- Distills agreeing rollouts via OPSD;
- Penalizes disagreeing rollouts with Grouped RL.
- Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks.
- In test-time training, TTPO improves Qwen3-1.7B from 38.0% to 45.2%.
- Gains of +25.2% to +36.4% in no-think mode.
- Demonstrates strong cross-task generalization.
Summary
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token.
The authors observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, they propose Test-Time Policy Optimization (TTPO), an asymmetric objective that:
Token-level selection further refines both branches: distillation downweights already-converged positions, while RL penalizes only confident mistakes. As a result, even with frequent pseudo-label errors, both updates remain well-grounded, and majority-vote routing provides progressively stricter self-supervision as the model improves.
Results
*Auto-collected on 2026-08-30.*