Summary
This arXiv paper (2608.20290) by Cheng Xu, Nan Yan, Liming Chen, and M-Tahar Kechadi audits claims of language model self-improvement by differencing noisy accuracy estimates across individual questions. The authors run three rounds of rank-32 LoRA self-training on Qwen3-8B, benchmarked against a frozen control passed through the identical pipeline. They identify seven measurement failure modes, each capable of reversing reported findings when the control is absent. Single-greedy-decode bookkeeping fabricates capability changes on an untrained model, largely due to inference batching artifacts, and an extension statistic distinguishing acquisition from sharpening yields a 0.280 ratio for this model. Natural threshold fixes fail under replication: the null hypothesis remains nonzero when estimated on frozen comparisons. Replacing thresholds with per-question exact tests against a pooled baseline under false discovery rate control detects no effects on any held-out replication. The audit finds that external distillation improves problems the base model rarely solves, while three self-training variants do not; regression shows this asymmetry is a byproduct of distillation's larger overall gains.
Paper Overview
Field: NLP
Authors: Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi
Published: 2026-08-22
arXiv: 2608.20290
Abstract (translated)
Whether language models self-improve is increasingly judged not by average accuracy but by which individual questions they gain and lose. Tracking these transitions means differencing two noisy estimates, leaving the analysis vulnerable to measurement artifacts. This paper audits three rounds of rank-32 LoRA self-training on Qwen3-8B against a frozen control passed through the identical pipeline, identifying seven measurement failures, each of which reverses the reported finding when its control is absent.
Key findings:
- Bookkeeping based on single-greedy-decode fabricates capability changes on an untrained model, driven largely by inference batching artifacts.
- An extension statistic that distinguishes acquisition from sharpening assigns this model a ratio of 0.280.
- Natural threshold fixes do not survive replication: when estimated on frozen comparisons, their null hypothesis remains nonzero.
- Replacing thresholds with per-question exact tests against a pooled baseline under false discovery rate control detects no effects on any held-out replication.
- The audit finds that external distillation improves problems the base model rarely solves, while three forms of self-training do not; regression rejects this asymmetry as a byproduct of distillation's larger overall gains.
---
*Auto-collected on 2026-08-22*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633823