Summary
This arXiv paper (2504.06255) by Qimin Zhong, Hao Liao, and Haiming Qin examines whether large language models develop coherent internal world models. While next-token prediction (NTP) provides only one-step-ahead supervision, multi-token prediction (MTP) shows promise for learning structured representations. The authors provide a theoretical analysis of MTP's gradient inductive bias, supported by experiments, showing that MTP induces representational contractivity through gradient coupling and promotes convergence toward internal belief states. However, they find that standard MTP often suffers from structural hallucinations: discrete token supervision encourages illegal shortcuts in latent space that violate environmental constraints. To address this, the paper proposes Latent Semantic Enhancement MTP (LSE-MTP), which anchors predictions to ground-truth hidden state trajectories. Experiments on synthetic graphs and real-world Manhattan taxi data show that LSE-MTP bridges the gap between discrete tokens and continuous state representations, improves representation alignment, reduces structural hallucinations, and increases robustness to perturbations.
Paper Overview
Research Area: NLP
Authors: Qimin Zhong, Hao Liao, Haiming Qin
Published: 2025-04-08
arXiv: 2504.06255
Abstract
Whether Large Language Models (LLMs) develop coherent internal world models remains a core debate. While conventional Next-Token Prediction (NTP) focuses on one-step-ahead supervision, Multi-Token Prediction (MTP) has shown promise in learning more structured representations.
In this work, the authors provide a theoretical perspective analyzing the gradient inductive bias of MTP, supported by empirical evidence, showing that MTP promotes the convergence toward internal belief states by inducing representational contractivity via gradient coupling.
However, the paper reveals that standard MTP often suffers from structural hallucinations: discrete token supervision encourages illegal shortcuts in latent space that violate environmental constraints. To address this, the authors propose Latent Semantic Enhancement MTP (LSE-MTP), which anchors predictions to ground-truth hidden state trajectories.
Findings
Experiments on synthetic graphs and real-world Manhattan taxi data demonstrate that LSE-MTP:
- Effectively bridges the gap between discrete tokens and continuous state representations
- Enhances representation alignment
- Reduces structural hallucinations
- Improves robustness against perturbations
---
*Auto-collected on 2026-04-09.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177169681