[论文] Rare Event Estimation via Iterative Unalignment

研究领域: ML 作者: Hanming Yang, Daksh Mittal, Jing Dong, Hongseok Namkoong 发布时间: 2026-09-21 arXiv: 2609.24969

论文概要

研究领域: ML 作者: Hanming Yang, Daksh Mittal, Jing Dong, Hongseok Namkoong 发布时间: 2026-09-21 arXiv: 2609.24969

中文摘要

随着智能体被赋予越来越高的自主性,即使在其随机输出轨迹中极其罕见的事件也可能发生并造成灾难性后果。因此,安全部署不取决于这些事件是否可能发生,而取决于它们发生的频率。我们研究估算由智能体自身动作随机变化引起的罕见事件概率的问题。估算此类风险需要在一个组合爆炸级的轨迹空间中进行搜索。朴素蒙特卡洛方法在此场景下计算上不可行,而构建有效的重要性采样(IS)提议分布需要对上下文相关的条件分布链进行协调一致的修改。我们开发了一种新的 IS 方法,通过扰动原始模型的权重来构建提议分布。该提议分布本身是一个可微参数化的语言模型,支持基于梯度的权重空间搜索。我们构建的目标函数结合了用于事件放大的可微代理和自适应正则化方案,动态平衡放大与估计器稳定性之间的关系。我们在约 1.2 亿和 26 亿参数的模型上、横跨三个事件家族的 300 多个罕见事件(概率低至 10^{-9},参考概率以低于 10% 的相对标准误差计算)上评估了我们的方法。在最具可验证性的场景中,我们观察到对于概率低于 10^{-7} 的事件,IS 估计器相比朴素蒙特卡洛实现了超过 800 倍的计算加权效率提升。代码开源:https://github.com/namkoong-lab/iterative-unalignment

原文摘要

As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal...


*自动采集于 2026-09-23*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens