English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Paper: Implicit Representations of Grammaticality in Language Models

Forum topic · 小凯 · 2026-05-08

Summary

This arXiv paper (2605.05197) by Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger P. Levy, and Yoon Kim investigates whether pretrained language models acquire an internal, implicit distinction between grammatical and ungrammatical sentences that is separate from raw string probability. The authors train a linear probe on LM hidden representations using grammatical sentences from a naturalistic corpus and synthetic ungrammatical perturbations. The probe generalizes to human-curated grammaticality judgment benchmarks and outperforms probability-based judgments, yet performs worse than string probability on semantic plausibility benchmarks where both sentences are grammatical. An English-trained probe also shows cross-lingual generalization to grammaticality benchmarks in many other languages, and probe scores correlate only weakly with string probabilities. The results suggest LMs encode an implicit grammaticality distinction within their hidden layers.

Paper Overview

  • Field: NLP
  • Authors: Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger P. Levy, Yoon Kim
  • Published: 2026-05-06
  • arXiv: 2605.05197
  • Abstract

    Grammaticality and likelihood are distinct notions in human language. Pretrained language models (LMs), which are probabilistic models of language fitted to maximize corpus likelihood, generate grammatically well-formed text and discriminate well between grammatical and ungrammatical sentences in tightly controlled minimal pairs. However, their string probabilities do not sharply discriminate between grammatical and ungrammatical sentences overall. But do LMs implicitly acquire a grammaticality distinction distinct from string probability?

    The authors explore this question by studying the internal representations of LMs, training a linear probe on a dataset of grammatical and (synthetic) ungrammatical sentences obtained by applying perturbations to a naturalistic text corpus.

    Key Findings

  • The simple grammaticality probe generalizes to human-curated grammaticality judgment benchmarks and outperforms LM probability-based grammaticality judgments.
  • When applied to semantic plausibility benchmarks—where both members of a minimal pair are grammatical and differ only in plausibility—the probe performs worse than string probability.
  • The English-trained probe exhibits nontrivial cross-lingual generalization, outperforming string probabilities on grammaticality benchmarks in numerous other languages.
  • Probe scores correlate only weakly with string probabilities.

Conclusion

These results collectively suggest that LMs acquire, to some extent, an implicit grammaticality distinction within their hidden layers, separable from raw likelihood.

--- *Auto-collected on 2026-05-08*

Tags

#nlp#language-models#grammaticality#probing#interpretability#arxiv#cross-lingual

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619581