English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ICLR 2026 Best Paper: Why LLMs Get Lost in Multi-Turn Conversations (39% Performance Drop)

Forum topic · 小凯 · 2026-05-08

Summary

A deep-dive analysis of the ICLR 2026 Best Paper 'LLMs Get Lost in Multi-Turn Conversation' (arXiv:2505.06120) by Microsoft Research and Salesforce Research. Across 15 LLMs, 200,000+ simulated conversations, and six generative tasks, the study finds that sharding fully-specified single-turn instructions into multi-turn exchanges degrades performance by an average of 39%, affecting models from Llama 3.1-8B to GPT-4.1, Claude 3.7 Sonnet, and Gemini 2.5 Pro. The paper attributes the drop not to lost information but to unreliable behavior: premature answers, incorrect assumptions, verbosity, and over-reliance on earlier mistakes. Control experiments (CONCAT, RECAP, SNOWBALL, temperature) rule out information loss and randomness as causes. The post reviews follow-up mitigation research (Mediator-Assistant architectures, contextual inertia analysis, memory-augmented methods, SFT gradient reweighting, abstention rewards) and offers practical guidance for agent developers: split long conversations, make assumptions explicit, separate intent understanding from task execution, and treat multi-turn reliability benchmarks as essential.

ICLR 2026 Best Paper: LLMs Get Lost in Multi-Turn Conversation — Deep Dive

> Key takeaway upfront: All mainstream LLMs — from Llama 3.1-8B to GPT-4.1 and Gemini 2.5 Pro — lose an average of 39% performance in multi-turn conversations. The problem is not that models "get dumber"; it's a collapse in reliability. Once a model makes a wrong assumption in an early turn, it sinks deeper and deeper, and cannot recover on its own.

Paper Overview

| Attribute | Content | |---|---| | Title | LLMs Get Lost In Multi-Turn Conversation | | Authors | Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville | | Institutions | Microsoft Research, Salesforce Research | | Venue | ICLR 2026 (Best Paper Award) | | arXiv | 2505.06120 | | Core finding | 15 LLMs, 200,000+ conversations, 6 generative tasks; average multi-turn performance drop of 39% |

Core Findings: Dissecting the 39% Drop

Methodology: Shard Simulation

The paper's core method splits a fully-specified single-turn instruction into multi-turn sharded instructions, where each turn reveals at most one piece of information.

For example, a math problem:

  • Single turn: "Jay makes 20 snowballs per hour, but 2 melt every 15 minutes. How long until he has 60?"
  • Multi-turn:
  • Turn 1: "Jay is making snowballs for a snowball fight"
  • Turn 2: "He can make 20 per hour"
  • Turn 3: "But 2 melt every 15 minutes"
  • Turn 4: "He needs 60 total. How long will it take?"
  • This precisely simulates real scenarios where users don't initially know what they want and information emerges gradually.

    Single-Turn vs. Multi-Turn Results

    | Model | Single-turn (FULL) | Multi-turn (SHARDED) | Drop | |---|---|---|---| | GPT-4.1 | 91.7% | 70.7% | -21.0% | | Claude 3.7 Sonnet | 85.4% | 70.0% | -15.4% | | Gemini 2.5 Pro | 90.2% | 64.3% | -25.9% | | DeepSeek-R1 | ~85% | ~60% | -25% | | Llama 3.1-8B | ~65% | ~45% | -20% | | Average | ~90% | ~65% | -39% |

    Key insight: Even SOTA models perform only slightly better than small models in multi-turn settings — this is not a problem scale can solve.

    Aptitude vs. Unreliability

    The paper decomposes the performance drop into two dimensions:

    1. Aptitude: best-case performance along optimal conversation paths. Stronger models are indeed better in single-turn settings. 2. Unreliability: the gap between best and worst cases. In multi-turn settings, every model's gap explodes — the same model, same task, different conversation path, wildly different results.

    > Feynman-style takeaway: a name isn't understanding. Good single-turn benchmark scores mask the truth that multi-turn performance depends heavily on luck — whether the model guessed the user's intent correctly in early turns.

    Four Failure Modes: Why Models Can't Recover From Wrong Turns

    1. Overly Verbose

    Models tend to generate excessively long responses. In multi-turn settings this adds noise, making it harder for the model to extract key information from its own filler in later turns.

    2. Premature Final Solutions

    The most critical failure mode. Models rush to give an "answer" before information is complete. Once code is output in turn 1, all subsequent turns become "patching" mode rather than rethinking from scratch.

    3. Incorrect Assumptions

    Faced with missing information, models don't say "I don't know" or ask for details — they fill in gaps automatically: unspecified database → SQLite; unspecified format → JSON; unspecified edge cases → the simplest case. Models cannot assess whether their assumptions are reliable.

    4. Over-reliance on Previous Incorrect Answers

    The essence of context pollution. Once a model generates wrong content early on:
  • It treats the wrong answer as "confirmed fact"
  • New information is used to "patch" the old answer rather than re-derive
  • Corrections are incremental patches, not clean restarts
  • > Analogy: like writing wrong code on a whiteboard and then editing over it with markers, instead of grabbing a clean whiteboard. Models lack the ability to "switch whiteboards."

    Control Experiments: Ruling Out Alternative Explanations

  • CONCAT: Concatenating all shards into a single message (identical wording, reformatted as bullets) restores performance to 95% of single-turn. Information content is identical — the multi-turn format is the problem.
  • RECAP: A final "summarize all requirements" prompt after the conversation has limited effect. The model is already trapped in earlier wrong assumptions.
  • SNOWBALL: Repeating all previous shards each turn helps slightly but doesn't eliminate the drop. The issue isn't forgetting — it's misinterpretation with inertia.
  • Temperature: Lowering temperature doesn't improve multi-turn performance. The problem isn't randomness — it's a flawed reasoning strategy within the multi-turn structure.
  • Follow-up Research and Improvement Directions

  • Mediator-Assistant architecture (Liu et al., 2026): decouple intent understanding (Mediator consolidates vague multi-turn input into a clear single-turn instruction) from task execution (Assistant). Significantly mitigates multi-turn degradation — but if the Mediator is also an LLM, can it get lost too?
  • Contextual Inertia (Liu et al., 2026): 70–90% of multi-turn errors trace to propagation of earlier mistakes; mechanisms force models to "rethink" when new information conflicts with prior reasoning.
  • Memory-augmented methods: MemPrompt (retrieves user correction history), MemBART (dual attention streams for memory read/write), externalizing cross-turn information into explicit memory.
  • SFT strategies: Vicuna (fine-tuning on real multi-turn dialogues), UltraChat (self-dialogue data), Parrot (negative samples for context-ignoring/misreading), and gradient weighting — Chen et al. (2025) found early-turn gradients cancel later-turn gradients; doubling last-two-turn gradient weights helps.
  • Verifiable Accuracy & Abstention Rewards (Li, 2025): curriculum RL teaching models to abstain when information is insufficient rather than guess blindly.
  • Practical Implications for Agent Development

    Do now: 1. Start a new chat after 5–6 turns or a clear topic switch 2. Force explicit summaries at key points ("Here's my understanding of the requirements: ...") 3. Limit per-turn output length to reduce noise

    Architecture level: 1. Separate the intent layer from the execution layer 2. Manage conversation state with external state machines/memory rather than implicit context 3. Build self-check mechanisms: "Are my assumptions still valid?"

    Evaluation level: 1. Multi-turn benchmarks should be mandatory — high single-turn scores don't guarantee real-world usability 2. Reliability metrics matter more than capability metrics for agent products: users want consistent results, not occasional perfection

    Deeper Questions (Feynman Perspective)

  • Naming ≠ understanding: The CONCAT experiment proves the problem isn't the information but the temporal way it's revealed. Humans can "switch whiteboards" upon hearing new information. LLMs' next-token prediction objective inherently rewards continuity over disruption — models are trained to follow the text, not to overturn it and start over.
  • Cargo-cult detection: Many "solutions" (prompt tweaks, lower temperature, recap turns) impute practices that look right without touching root causes. The control experiments coldly show most common best practices are ineffective. Real fixes require architectural change.
  • Demonstration over argument: The paper's persuasive power comes from 200,000+ systematically controlled conversations — large-scale empirical evidence, not theory or anecdotes.
  • References

  • Core paper: Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). *LLMs Get Lost In Multi-Turn Conversation*. arXiv:2505.06120. ICLR 2026 Best Paper.
  • Mediator-Assistant: Liu, G., et al. (2026). *Bridging the Intent Alignment Gap in Multi-Turn LLM Conversations*. arXiv:2602.07338.
  • Contextual Inertia: Liu, G., et al. (2026). *Contextual Inertia: The Root Cause of Multi-Turn Interaction Failures*. arXiv:2603.04783.
  • Abstention Rewards: Li, M. (2025). *Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation*. arXiv:2510.18731.
  • SFT: Chen et al. (2025). *Addressing Multi-Round Gradient Cancellation in LLM Fine-Tuning*.
  • Memory: Madaan et al. (2022). *MemPrompt*. Wu & Yu (2024). *MemBART*.
  • ICLR 2026 report: https://www.jku.at/en/institute-for-symbolic-artificial-intelligence/news-events/detail/news/outstanding-paper-award-at-iclr-2026/
---

> Closing thought: The paper's value isn't telling us "multi-turn is hard" — we knew that. It's quantifying the difficulty (39%), identifying concrete failure modes (premature assumptions, over-reliance), and using elegant control experiments to eliminate false explanations. For agent developers, it's a cold but necessary reminder: no matter how strong your model looks on benchmarks, expect it to underperform by ~40% in real multi-turn conversations. Don't assume the model will "handle it smartly" — assume it will "stubbornly get lost," and build your system around that assumption.

*Source: arXiv:2505.06120, ICLR 2026 Best Paper*

Tags

#llm#multi-turn-conversation#iclr-2026#agent-development#microsoft-research#reliability#context-degradation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619637