Paper Overview
Field: NLP Authors: Enric Boix-Adsera, Benedict Tessler Published: 2026-08-17 arXiv: 2608.16834
Abstract (Original)
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.
Key Points
- Model hypnosis: individually weak, seemingly irrelevant prompt cues can be systematically combined to strongly steer model behavior.
- Generality: the effect spans model families and scales, including frontier reasoning models.
- Transferability: hypnotic prompts can transfer between different models.
- Inconspicuous control: manipulation relies on subtle textual choices such as paraphrases and typos.
- Implications: new challenges and avenues for AI safety, and a major hurdle for AI interpretability.