[论文] Online Safety Monitoring for LLMs

研究领域: NLP 作者: Mona Schirmer, Metod Jazbec, Alexander Timans 发布时间: 2026-07-04 arXiv: 2507.00479

论文概要

研究领域: NLP 作者: Mona Schirmer, Metod Jazbec, Alexander Timans 发布时间: 2026-07-04 arXiv: 2507.00479

中文摘要

尽管经过对齐训练,大语言模型在部署时仍然容易产生不安全输出。因此,在线监控输出并在安全不再有保障时发出警报至关重要。我们研究了一种简单的实时监控器,通过阈值化将外部模型的验证器信号转化为警报决策,阈值通过风险控制校准。在数学推理和红队数据集上的实验中,我们表明这种简单设计与基于序贯假设检验的更先进监控器相比具有竞争力。

原文摘要

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.


*自动采集于 2026-07-04*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens