Paper Overview
Research Area: NLP Authors: Cheng Qian, Wenting Zhao, Liangwei Yang arXiv: 2508.03418
Summary (translated from Chinese)
Recent research on distillation typically transfers large-model capabilities to smaller models by updating the smaller model's parameters via teacher forcing, online distillation, and related training-time methods. This paper asks whether such transfer can happen at test time. The authors study strong-to-weak scaffolding: can a stronger builder model construct inference-time harnesses that help a weaker target model solve tasks more reliably, without any parameter updates? Using four representative theory-of-mind benchmarks, each builder model uses 5% of the data as a validation set and iteratively refines its harness over multiple rounds; the final harness is then evaluated on the full test set.
Empirically, this test-time capability transfer is highly effective: target model performance nearly doubles on average from 0.49 to 0.91. Analysis shows the gains mainly come from offloading unstable model reasoning to deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than from encouraging the target model to reason more broadly or sample more widely. The authors further find that builder models' reasoning effort monotonically improves harness quality; plateau effects are modest relative to the builder's own capability, and weaker target models gain the most.
These results indicate that inference-time harness design is an important complement to traditional training-time distillation, enabling strong models to transfer cognitive structure to weak models without retraining.
Original Abstract (excerpt)
> Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters...