English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Phantom Gains: Auditing Self-Improvement Against a Measured Null (arXiv 2608.20290)

Forum topic · 小凯 · 2026-08-22

Summary

This paper, 'Phantom Gains: Auditing Self-Improvement Against a Measured Null' by Cheng Xu, Nan Yan, Liming Chen, and M-Tahar Kechadi (arXiv:2608.20290, released 2026-08-22, NLP), audits claims of language-model self-improvement. Rather than judging improvement by average accuracy, it tracks which individual questions a model gains or loses, differencing two noisy estimates—a procedure vulnerable to measurement artifacts. The authors audit three rounds of rank-32 LoRA self-training on Qwen3-8B against a frozen control run through the identical pipeline, identifying seven measurement failures, each capable of reversing the reported finding when the control is absent. Single-pass greedy decoding bookkeeping fabricates capability changes on an untrained model, largely inference-batching artifacts; an extension statistic distinguishing acquisition from sharpening yields a ratio of 0.280. Natural threshold fixes fail replication because their null remains nonzero under the frozen comparison. Replacing them with per-question exact tests against a pooled baseline under false discovery rate control detects no effects on any held-out replication. The audit finds external distillation improves problems the base model rarely solves, while three forms of self-training do not; regression attributes this asymmetry as a byproduct of distillation's larger overall gains.

Paper Overview

Field: NLP Authors: Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi Released: 2026-08-22 arXiv: 2608.20290

Abstract (English translation)

Whether language models self-improve is increasingly judged not by average accuracy, but by which individual questions they gain and lose. Tracking these transitions means differencing two noisy estimates, making it vulnerable to measurement artifacts.

This paper audits three rounds of rank-32 LoRA self-training on Qwen3-8B against a frozen control run through the identical pipeline, identifying seven measurement failures, each of which reverses the reported finding in the absence of its control.

  • A ledger based on single-pass greedy decoding manufactures capability changes on an untrained model, driven largely by inference-batching artifacts.
  • An extension statistic distinguishing acquisition from sharpening assigns a ratio of 0.280 to the model.
  • Natural threshold fixes fail replication: estimated against the frozen comparison, their null hypothesis remains nonzero.
  • Replacing them with per-question exact tests against a pooled baseline under false discovery rate control detects no results on any held-out replication.
The audit finds that external distillation improves problems the base model rarely solves, while three forms of self-training do not; regression rejects this asymmetry as a byproduct of distillation's larger overall gains.

---

*Auto-collected on 2026-08-22*

Tags

#nlp#arxiv#self-improvement#lora#qwen3-8b#measurement-audit#distillation#evaluation-methodology

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633802