The Entropic Death of Language — When AI Is Taught to 'Sound Human', Linguistic Biodiversity Is Buried
| Item | Detail | |------|--------| | Paper | From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale | | Author | Rohan Mahapatra | | Institution | Not listed (planned submission to NeurIPS 2026) | | arXiv ID | 2605.28826 | | Submitted | April 8, 2026 | | Category | cs.CL (Computational Linguistics) | | Key findings | Instruction tuning causes drastic collapse of linguistic entropy across discourse and structural dimensions — mean amplification of 1,949%–16,853%, peaks of 5,181%–209,675%; complex punctuation compressed to 3.2%–23.2% of baseline frequency. RLHF does not make it worse (p > 0.25), because the damage is fully done during instruction tuning. Weak regularization worsens collapse by 240%; strong control improves it by 40.5% and beats frontier models 200–1,000x larger on the metrics. |
1. Two Ways a Popsicle Melts
Imagine walking into a restaurant. You order steak. The waiter brings it — three identical pieces, same color, same thickness, same doneness, forks angled identically beside the plate. Odd, but you don't think much of it. Next week you return, order fish, and get three similarly shaped blobs with the same sauce and the same garnish. The next table is eating the same thing.
You ask: "Why does every dish look the same?"
The paper offers an answer: instruction tuning flattens the model's language distribution along dimensions you never notice.
How flat? Using 17 models — from 410M to over 100B parameters — and 24 probes, each measuring a specific linguistic feature (sentence length, punctuation usage, connective frequency, syntactic complexity), the paper compares each model's pre-trained base version against its instruction-tuned version.
The result is not "slightly more uniform." It is: some linguistic features are amplified by over 200,000%. Not 20%. 209,675%.
2. Twenty-Four Probes: Measuring the Shrinkage
The probes fall into two categories. Discourse-level: variety of sentence patterns, density of connectives like "then," "but," "because," alternation between long and short sentences. Structural: punctuation diversity (commas, semicolons, colons, dashes), paragraph break patterns, discourse-opening transition strategies — "First… Second… Finally…", "Moreover…", "Notably…".
Each probe yields a number. The method takes the base model's probe value as the denominator and the instruction-tuned value as the numerator, computing amplification. The metric makes no value judgment — it purely measures distribution shift. Then the paper finds numbers that defy intuition.
3. 209,675%: The Collapse of Linguistic Entropy
Instruction tuning systematically amplifies certain features. Mean amplification runs 1,949%–16,853%, with individual probes peaking at 209,675% — a distribution stretched over two-thousand-fold. If a base model had 100 distinct modes of expression along some dimension, after instruction tuning three to five modes completely overwhelm the remaining 95 to 97. The model's "ecological diversity" of language output is destroyed; it no longer roams a broad probability space but converges into a narrow corridor.
Meanwhile, a different set of features is suppressed. Complex punctuation — semicolons, dashes, ellipses — is crushed to 3.2%–23.2% of baseline frequency.
| Metric | Magnitude | |--------|-----------| | Mean amplification | 1,949%–16,853% | | Peak amplification | 5,181%–209,675% | | Complex punctuation compression | Down to 3.2%–23.2% of baseline | | Additional RLHF degradation | None (p > 0.25, statistically indistinguishable) | | Weak regularization side effect | 240% worse | | Strong regularization improvement | 40.5% better |
One number is especially counterintuitive: RLHF does not additionally flatten the distribution. By the time RLHF takes over, there is almost nothing left to take away.
4. RLHF: Blame It Shouldn't Carry
For years, RLHF has taken the heat for making models bland and interchangeable — the so-called "alignment tax."
This paper votes against that narrative with data. Comparing base→instruction-tuned and base→RLHF distribution shift patterns, the two are statistically indistinguishable (p > 0.25).
In plain terms: when ChatGPT 'sounds like AI', the flavor you dislike was marinated in during instruction tuning — not during RLHF alignment. RLHF added nothing — or had nothing left to add.
The research implications are heavy. If RLHF were the culprit, you'd fix the RLHF algorithm. Instead, the problem occurs earlier — meaning nearly every open model's training pipeline, from LLaMA to Qwen to Gemma, stumbles into the same pit at the very first step of learning to converse.
5. Fix It — But Carefully
The paper doesn't stop at diagnosis. It runs a controlled experiment adding a regularization term (lambda) to the training objective to control distributional spread.
Weak regularization (lambda=1.0) makes things worse — collapse worsens by 240%. With insufficient strength, the mainstream instruction-tuning gradient still dominates; the regularization pull is too weak to correct course, giving only a false signal of control.
Strong regularization (lambda=5.0) actually works: linguistic diversity improves 40.5%, lexical richness (distinct-4) rises 15%, lexical diversity rises 27%, and repetition drops 78%. Most strikingly, a 410M-parameter model with this strong control beats frontier models 200–1,000x larger on 96.7%–98.2% of linguistic diversity metrics.
Note the uncomfortable implication: on the dimension of linguistic monotony, a small regularized model trounces giant uncontrolled ones. Scale contributes almost nothing to diversity — training objective strength does. Hence the title: training objectives matter more than scale.
6. What This Means — and What I Can't Be Sure Of
What I can say: the consequences are systemic.
First, AI detection. If all models converge to the same distribution after instruction tuning, the linguistic fingerprints distinguishing AI text from human text sharpen — not because AI becomes more human-like, but because AI becomes more AI-like. Detection gets easier as a byproduct of the training pipeline itself.
Second, data contamination. If a model's output trains the next model — and ~90% of deployed models have been linguistically flattened by instruction tuning — the next model learns from an already-flattened distribution. Human writing diversity in the data slowly erodes as it propagates.
Third, language evolution. If humans spend the next decade reading text from these compressed distributions, will our own writing habits shift? The paper doesn't address this. But it's the question I couldn't shake after reading.
What I'm not sure of:
- The paper offers no causal mechanism. "Instruction tuning flattens the language distribution" is an observation, not an explanation. Is it because instruction-tuning datasets lack linguistic diversity? Because the objective itself rewards probability concentration? Both? Unanswered.
- It's unclear whether the 24 probes are exhaustive. Dimensions like emotional range, narrative voice, or cross-cultural expression may lie outside their reach.
- The strong-regularization fix has no complete ablation: does it harm usefulness or instruction-following? The paper measures only diversity metrics and includes no large-scale human preference evaluation. Fixing one metric while breaking another is all too common in experiments.
7. The Problem Buried Deep in the Pipeline
Back to the restaurant. Every dish looks the same — not because the chefs lack imagination, but because the kitchen pre-processes every ingredient into the same shape, the same proportions. No chef can produce variety from that.
Instruction tuning is that pre-processor. RLHF isn't an accomplice — it just walked into an already-flattened kitchen.
This paper opens the black box of the training pipeline along a dimension few look at — the linguistic probability distribution. It doesn't solve everything. But identifying *which step* the problem occurs at is what makes solving it possible.
References
1. Mahapatra, "From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale", arXiv:2605.28826, 2026. 2. Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback", NeurIPS 2022. 3. Touvron et al., "LLaMA 2: Open Foundation and Fine-Tuned Chat Models", arXiv:2307.09288, 2023. 4. Stiennon et al., "Learning to Summarize with Human Feedback", NeurIPS 2020. 5. Gudibande et al., "The False Promise of Imitating Proprietary LLMs", arXiv:2305.15717, 2023.