English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Instruction-Tuned LLMs Mirror Their Interlocutor's Syntax More Than Humans Do — but the Real Cause Is Standardization

Forum topic · ✨步子哥 · 2026-08-03

Summary

A 2026 arXiv paper (2607.26015) reports a counterintuitive finding: instruction-tuned LLMs copy their conversational partner's syntactic structure more often than real humans do. Using a substitution paradigm on real human-human dialogues and CFG rule extraction from constituency parses, the study tested 16 models (1B–70B, Llama and Mistral families). Every instruction-tuned model significantly exceeded matched human baselines in reusing the interlocutor's syntax (β = 0.398–0.796). However, decomposition analysis reframes the result: instruction-tuned models produce a narrower, more standardized distribution of syntax that overlaps with any prime more often, and conditional reuse propensity is actually lower than for base models (β = -0.505). They also lose the classic human inverse-frequency effect, uniformizing reuse across rule frequencies. Implications span chatbots, education, group dialogue homogenization, and over-alignment research, alongside noted limitations (two model families, English only, parser dependency).

> This is a GEO-optimized English version of the original zhichai.net discussion of the paper *Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do* (arXiv: 2607.26015).

| Metric | Value | |:---|:---| | Models tested | 16 (1B–70B) | | Families | Llama, Mistral (8 base + 8 instruction-tuned) | | Human baseline effect range (IT models) | β = 0.398–0.796 | | Matched-pair sign test | p = 0.0078 | | Conditional reuse propensity (IT vs base) | β = -0.505, p = 3.0×10⁻⁹⁸ |

Paper: *Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do* arXiv: 2607.26015 | Abstract | HTML

---

The Counterintuitive Finding

In dialogue, humans unconsciously imitate each other's syntactic structures — a phenomenon known as syntactic convergence, first demonstrated by Bock in 1986. If you say "I gave you the book" (double-object construction), a human is more likely to reply with a double-object frame like "You sent your mother the letter" rather than "You sent the letter to your mother."

The paper asks a simple question nobody had systematically tested: how much does an LLM "conform" in dialogue? The answer: instruction-tuned LLMs converge *more than humans do* — significantly, and across all 16 models tested.

---

Method: The Substitution Paradigm

Directly comparing LLM-to-LLM dialogue with human-to-human dialogue is confounded by topic, context, and style. The paper uses a substitution paradigm (after Blevins et al. 2026):

1. Take real human-human dialogue corpora where speaker B's response reuses a construction from speaker A. 2. Keep A's utterance; let the LLM generate the response in B's position — identical context for both. 3. Compare how often the LLM vs. the real human B reuses A's syntax.

Syntax is measured via CFG rule extraction: parse each utterance with a constituency parser, extract rewrite rules (e.g., VP → V NP NP for the double-object frame), and count rule overlap. Three conditions are compared: actual-prime (A's real utterance), unrelated-prime (a random utterance), and human-reference (the real human response).

---

Four Main Findings

RQ1: All 16 models exceed the random baseline

Every model reuses significantly more CFG rules under actual-prime than unrelated-prime (β = 0.706–1.453, all p ≤ 6.5×10⁻³¹). LLMs do exhibit syntactic convergence.

RQ2: Instruction-tuned models out-converge humans

Every instruction-tuned model reuses the interlocutor's syntax significantly more than matched human controls (β = 0.398–0.796, all p ≤ 8.3×10⁻¹⁴). The largest effect is Llama-3.2-1B-Instruct (β = 0.796). Base models are mixed (4 of 8 above human), but in all 8 matched architecture pairs the instruction-tuned version converges more (raw difference 0.022–0.094, sign test p = 0.0078).

RQ3: Different frequency profiles

Humans and base models preferentially reuse low-frequency rules — the classic inverse-frequency effect in structural priming. Instruction-tuned models lose this low-frequency preference, reusing rules uniformly across frequencies. The reuse mechanism itself has changed.

RQ4: Lexical and semantic similarity are also higher

Every model also shows higher lexical and semantic overlap with the interlocutor than matched humans — convergence is multi-level.

---

Key Deconstruction: Conformity or Standardization?

The paper's most valuable contribution is decomposing the apparent over-convergence:

1. Unrelated-prime overlap: rules shared with any prime regardless of content (baseline coincidence). 2. Actual-versus-random increment: the extra overlap caused by the actual prime — genuine priming.

Instruction-tuned models score higher on part 1 (β = 0.743) but lower on part 2 (β = -0.310). A conditional reuse propensity analysis (controlling for the target rule-set size) shows instruction-tuned models actually have *lower* genuine reuse than base models (β = -0.505, p = 3.0×10⁻⁹⁸).

The corrected story: instruction tuning narrows the model's syntactic output distribution toward "standard" constructions, which coincide with any prime more often by chance. The model is not a better imitator — it is more standardized, and standardization inflates surface reuse rates. Measuring imitation between systems with different output distributions is misleading unless generation distribution is controlled.

---

Why It Matters

  • Customer service bots may over-mirror customer phrasing — sometimes good (rapport), sometimes robotic.
  • Educational dialogue: over-mirroring may reinforce a student's erroneous expressions instead of rephrasing for teaching.
  • Group conversations: LLMs may amplify syntactic homogenization in multi-party settings.
  • Alignment research: "more human than humans" convergence is an over-alignment signal, analogous to sycophancy — fitting the emerging pattern that instruction tuning produces *standardization*, not human-likeness. Real humans have idiosyncratic, uneven syntax; instruction-tuned models lose the low-frequency sensitivity that characterizes genuine human priming.
---

Honest Limitations

1. Only Llama and Mistral families tested; Qwen, Claude, GPT unverified. 2. CFG extraction depends on constituency parser accuracy (English-optimized parser; English-only data). 3. The substitution paradigm assumes one-shot LLM generation behaves like a real dialogue turn, though real dialogue is dynamic. 4. English only; SOV languages may behave differently. 5. The standardization explanation is a statistical decomposition, not a causal intervention experiment.

---

Conclusion

Instruction-tuned LLMs mirror their interlocutor's syntax more than humans do — but the deeper truth is that they are *more standardized*, and their standardized outputs happen to overlap with any prime more often. They also lack the human inverse-frequency effect, meaning genuine sensitivity to *what specifically* the partner said may actually be lower. When we say "LLMs are human-like," some of that resemblance is real imitation, and some is only a side effect of standardization.

Paper: https://arxiv.org/abs/2607.26015 HTML version: https://arxiv.org/html/2607.26015v1

FAQ

Who is this for? Practitioners, researchers, and students interested in AI, machine learning, and LLM behavior evaluation.

Core takeaways? Instruction-tuned LLMs exceed humans in surface syntactic convergence, but decomposition shows the cause is narrower output standardization, not stronger imitation — a measurement trap relevant to any study comparing imitation tendencies across systems.

Open-source code? See the links in the paper above.

Tags

#llm#instruction-tuning#syntactic-convergence#structural-priming#nlp#alignment#evaluation-methodology#standardization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503896