A deep-dive report on the ICLR 2026 Best Paper 'LLMs Get Lost in Multi-Turn Conversation' (arXiv:2505.06120) by Laban et al. from Microsoft Research and Salesforce Research. Using a Sharded Simulation methodology, the study splits single-turn instructions into progressive multi-turn dialogues and tests 15 frontier models across 6 task domains over 200,000+ simulated conversations. Average performance drops from ~90% (single-turn) to ~65% (multi-turn), a ~39% degradation starting as early as turn two. Crucially, the drop stems not from reduced aptitude but a massive spike in unreliability (performance variance), and lowering temperature to 0.0 barely helps. The paper identifies four failure modes: premature answering, answer bloat (over-reliance on own mistakes), loss of middle turns, and verbosity breeding assumptions. Simple interventions (RECAP, SNOWBALL) only partially mitigate. The report also reviews follow-up work including Mediator-Assistant architectures, RLAAR abstention-reward RL training, and memory-augmented approaches, plus practical guidance for LLM builders, agent developers, and end users.
ICLR 2026 Best Paper Deep Dive: LLMs Get Lost in Multi-Turn Conversation
Summary
A deep-dive report on the ICLR 2026 Best Paper 'LLMs Get Lost in Multi-Turn Conversation' (arXiv:2505.06120) by Laban et al. from Microsoft Research and Salesforce Research. Using a Sharded Simulation methodology, the study splits single-turn instructions into progressive multi-turn dialogues and tests 15 frontier models across 6 task domains over 200,000+ simulated conversations. Average performance drops from ~90% (single-turn) to ~65% (multi-turn), a ~39% degradation starting as early as turn two. Crucially, the drop stems not from reduced aptitude but a massive spike in unreliability (performance variance), and lowering temperature to 0.0 barely helps. The paper identifies four failure modes: premature answering, answer bloat (over-reliance on own mistakes), loss of middle turns, and verbosity breeding assumptions. Simple interventions (RECAP, SNOWBALL) only partially mitigate. The report also reviews follow-up work including Mediator-Assistant architectures, RLAAR abstention-reward RL training, and memory-augmented approaches, plus practical guidance for LLM builders, agent developers, and end users.
This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619639