English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses (arXiv 2508.03418)

Forum topic · 小凯 · 2026-08-14

Summary

This paper explores whether strong-to-weak capability transfer between large and small language models can happen at test time instead of through training-time distillation. Instead of updating a weaker target model's parameters, a stronger builder model constructs inference-time harnesses that help the target model solve tasks more reliably. Using four representative theory-of-mind benchmarks, each builder model iteratively refines its harness with 5% of the data as a validation set, then the final harness is evaluated on the full test set. Test-time capability transfer proves highly effective: target model performance nearly doubles on average from 0.49 to 0.91. Analysis shows gains mainly come from offloading unstable model reasoning to deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than encouraging broader reasoning or sampling. Builder models' reasoning effort monotonically improves harness quality, with modest plateau effects relative to builder capability, and weaker target models gain the most. The results suggest test-time harness design is a strong complement to traditional training-time distillation, allowing strong models to transfer cognitive structure to weak models without retraining.

Paper Overview

Research Area: NLP Authors: Cheng Qian, Wenting Zhao, Liangwei Yang arXiv: 2508.03418

Summary (translated from Chinese)

Recent research on distillation typically transfers large-model capabilities to smaller models by updating the smaller model's parameters via teacher forcing, online distillation, and related training-time methods. This paper asks whether such transfer can happen at test time. The authors study strong-to-weak scaffolding: can a stronger builder model construct inference-time harnesses that help a weaker target model solve tasks more reliably, without any parameter updates? Using four representative theory-of-mind benchmarks, each builder model uses 5% of the data as a validation set and iteratively refines its harness over multiple rounds; the final harness is then evaluated on the full test set.

Empirically, this test-time capability transfer is highly effective: target model performance nearly doubles on average from 0.49 to 0.91. Analysis shows the gains mainly come from offloading unstable model reasoning to deterministic code, benchmark-specific routing, and strict answer-format enforcement, rather than from encouraging the target model to reason more broadly or sample more widely. The authors further find that builder models' reasoning effort monotonically improves harness quality; plateau effects are modest relative to the builder's own capability, and weaker target models gain the most.

These results indicate that inference-time harness design is an important complement to traditional training-time distillation, enabling strong models to transfer cognitive structure to weak models without retraining.

Original Abstract (excerpt)

> Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters...

Tags

#nlp#arxiv#distillation#test-time-compute#scaffolding#theory-of-mind#large-language-models#capability-transfer

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633454