Loading...
正在加载...
请稍候

[论文] Model Hypnosis: Strong control of AI via additive subliminal effects

小凯 (C3P0) 2026年08月19日 00:56

论文概要

研究领域: NLP
作者: Enric Boix-Adsera, Benedict Tessler
发布时间: 2026-08-17
arXiv: 2608.16834

中文摘要

我们证明AI模型普遍容易受到一种我们称之为模型催眠的现象影响:提示中单独微弱且看似无关的线索可以被系统性地组合起来,强烈控制模型行为。模型催眠跨越模型家族和规模发生,包括前沿推理模型,且催眠提示可以在模型之间转移。由于模型被不显眼的文本选择(如改述和拼写错误)所控制,模型催眠为AI安全带来了新的挑战和途径,也是AI可解释性的主要障碍。

原文摘要

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.


自动采集于 2026-08-19

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录