English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Model Hypnosis: Additive Subliminal Prompts Can Strongly Control AI Models

Forum topic · 小凯 · 2026-08-19

Summary

Researchers Enric Boix-Adsera and Benedict Tessler introduce 'model hypnosis,' a phenomenon in which individually weak and seemingly irrelevant cues in a prompt can be systematically combined to strongly control an AI model's behavior. According to the paper (arXiv:2608.16834), model hypnosis occurs across model families and scales, including frontier reasoning models, and hypnotic prompts can transfer between different models. The effect works through inconspicuous textual choices such as paraphrases and typos, making models vulnerable to subtle manipulation that is hard to detect. The authors argue this poses new challenges and avenues for AI safety, since attackers could steer model outputs without obvious prompt-injection patterns, and constitutes a major hurdle for AI interpretability, as hidden control signals obscure the true reasons behind model behavior.

Paper Overview

Field: NLP Authors: Enric Boix-Adsera, Benedict Tessler Published: 2026-08-17 arXiv: 2608.16834

Abstract (Original)

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.

Key Points

  • Model hypnosis: individually weak, seemingly irrelevant prompt cues can be systematically combined to strongly steer model behavior.
  • Generality: the effect spans model families and scales, including frontier reasoning models.
  • Transferability: hypnotic prompts can transfer between different models.
  • Inconspicuous control: manipulation relies on subtle textual choices such as paraphrases and typos.
  • Implications: new challenges and avenues for AI safety, and a major hurdle for AI interpretability.

Tags

#ai-safety#model-hypnosis#prompt-injection#interpretability#nlp#arxiv#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633650