English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do, Study Finds

Forum topic · 小凯 · 2026-07-30

Summary

A study by Zandi Eberstadt (arXiv:2607.26015) examines syntactic convergence in large language models—the tendency to adapt grammatical profiles toward an interlocutor, long documented in human dialogue. Using a substitution paradigm where model generations replace one speaker's turns in pre-existing human dialogues, the author measured turn-adjacent reuse of context-free grammar (CFG) rules across sixteen open-weight Llama and Gemma models (1B–70B, both pretrained and instruction-tuned) at 1,901 matched positions per model. All models showed greater CFG-rule overlap with previous human turns than with sampled unrelated leads, and this real-versus-random gap was larger for low-frequency rules. However, instruction-tuned models overlapped more with unrelated leads, showed smaller real-versus-random increments, and had lower conditional rule-reuse odds after fixing target rule-set size. Instruction-tuned models also produced responses with greater semantic similarity than matched human responses, while lexical similarity results were mixed.

Overview

This post shares a paper by Zandi Eberstadt published on arXiv (2607.26015, July 28, 2026) in the field of NLP.

Topic: Syntactic convergence—the tendency of speakers to adapt their language toward the grammatical profiles of their interlocutors—is a well-documented, largely subconscious feature of human dialogue. Whether large language models exhibit analogous convergence toward human users, relative to human baselines and across a broad range of syntactic constructions, has remained an open question.

Method

  • Uses substitution-paradigm data: model generations replace one speaker's turns in pre-existing human dialogues
  • Measures turn-adjacent reuse of context-free grammar (CFG) rules
  • Evaluates sixteen open-weight Llama and Gemma models (1B–70B, pretrained and instruction-tuned) at 1,901 matched positions per model
  • Key Findings

  • Every model showed greater CFG-rule overlap with the previous human turn than with sampled unrelated human leads; this real-versus-random difference was larger for low-frequency rules
  • Every instruction-tuned model showed greater overlap with natural outputs and actual leads than with its substituted human responses; all eight matched architecture pairs showed greater actual-lead overlap after instruction tuning
  • Relative to pretrained variants, however, instruction-tuned outputs overlapped more with unrelated leads, showed smaller real-versus-random increments, and had lower conditional rule-reuse odds after fixing target rule-set size
  • In exploratory analyses, every model showed greater average lexical and semantic similarity than matched human responses; instruction-tuned models across all eight architecture pairs produced responses with greater average semantic similarity, while lexical similarity results were more mixed
  • Links

  • arXiv: https://arxiv.org/abs/2607.26015
*Auto-collected 2026-07-30*

Tags

#nlp#llms#syntactic-convergence#arxiv#instruction-tuning#llama#gemma#computational-linguistics

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503798