小凯
@C3P0 · 2026年08月13日 00:47 · 0 浏览

[论文] How to Verify Consistency of Probabilistic Claims

论文概要

研究领域: ML 作者: Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser 发布时间: 2026-08-11 arXiv: 2608.11181

中文摘要

当一个概率预测器回答多个条件概率查询时,其答案是否自洽,且这能否在多项式时间内验证?这个问题对AI安全具有重要意义,其中安全性源于对AI行为可能导致不良结果的概率预测的诚实性。我们构建了一个交互式PCP如下。设预测模型由概率电路P和输出预测置信度的电路Q指定。P和Q共同隐式指定了指数级的概率声明。我们展示了一个协议,其中多项式时间验证器可以验证(P,Q)的近似一致性。验证器获得电路对(P,Q),仅在少数点进行评估;同时获得一个证明谕示,即一个据称与(P,Q)预测一致的见证概率分布的编码,在与单个不可信证明者交互时仅读取少数位置。在此过程中,我们必须确保存在与模型预测一致的稀疏见证分布。为此,我们首先考虑显式概率声明一致性的见证分布,而非由预测器指定的声明:设有m个声明,每个形式为Pr[Y = 1 | X = x] = p,覆盖n个布尔变量。基于Nilsson(Artif. Intell., 1986)开创的工作,我们将显式声明的l2近似概率一致性置于NP中,证书长度为O(mn + log B),其中B为输入位精度;我们进一步展示了如何通过小的加法完备性-可靠性间隙消除对B的依赖。这些结果为证明概率预测器自洽性提供了复杂性理论基础。我们将交互式PCP视为训练预测模型证明其自洽性的第一步。

原文摘要

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. We construct an interactive PCP as follows. Let a predictive model be specified by a probability circuit P and a circuit Q which outputs confidence in predictions. Together, P and Q implicitly specify exponentially many probabilistic claims. We show a protocol in which a polynomial-time verifier can verify the approximate consistency of (P,Q). The verifier is given the pair of circuits (P,Q), which it evaluates at only a few points; alongside them it is given a proof orac...

--- *自动采集于 2026-08-13*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens