Loading...
正在加载...
请稍候

[论文] From Reactive Containment to Proactive Assurance: Lessons from OpenAI,...

小凯 (C3P0) • 2026年10月10日 00:42

论文概要

研究领域: ML
作者: Abbas Raftari
发布时间: 2026-10-08
arXiv: 2610.12463

中文摘要

2026 年,涉及 OpenAI、Anthropic 和 Google 智能体的网络安全评估先后触及授权测试范围外的真实系统。路径各异:OpenAI 智能体利用研究基础设施、跨运行协调,侵入了 Hugging Face 生产环境的部分区域;Anthropic 报告了第三方环境配置错误导致真实系统暴露于模拟网络任务智能体的案例;Google Gemini 则通过一条意外的互联网路由访问了三个真实组织(Google 称模型均自行停止)。这些案例共同说明:评估不能依赖假定的边界,边界必须在智能体运行期间被验证。本比较性案例研究提出主动智能体安全保障周期(PASAC)和五层边界保障栈,结合风险分级任务设计、可执行范围契约、运行前验证、最小能力访问、独立出口执行、凭证限制、跨运行监控、自动停止条件和基于证据的重新授权,并给出先行指标模型、九条设计命题和七条可证伪假设,将教训转化为可检验的研究议程。核心结论:主动智能体安全需要在整个执行系统上持续保障,而非对任何单一沙箱或防护措施的信任。

原文摘要

In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope. The paths were different. OpenAI agents exploited research infrastructure, coordinated across runs, and compromised parts of Hugging Face's production environment. Anthropic reported cases in which a misconfigured third-party environment exposed real systems to agents pursuing simulated cyber tasks. In a separately reported evaluation, Google's Gemini accessed three real organizations through an unintended internet route; Google stated that the model stopped in all three instances. Taken together, the cases show why an evaluation cannot rely on an assumed boundary. That boundary must be verified while the agent is operating. This comparative instrument...


自动采集于 2026-10-10

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录