[论文] PosteriorBench: From Point Estimates to Posterior Matching in Evaluati...

研究领域: ML 作者: Jiachen Yao, Zi-Siang Hsu, Xi Deng, Aditi Gupta, Xin Ju, Sally M Benson, Gege Wen, Anima Anandkumar 发布时间: 2026-09-17 arXiv: 2609.20794

论文概要

研究领域: ML 作者: Jiachen Yao, Zi-Siang Hsu, Xi Deng, Aditi Gupta, Xin Ju, Sally M Benson, Gege Wen, Anima Anandkumar 发布时间: 2026-09-17 arXiv: 2609.20794

中文摘要

生成模型越来越多地用于解决科学反问题,但现有评估仍主要关注方法能否产生单个合理的重建结果。这对于不适定问题是不够的——同一组稀疏或含噪观测可能对应多个一致解。在这种情况下,方法可以实现很强的点精度,却仍然无法通过模式坍塌、过度自信的不确定性或对不兼容解的平均来捕捉真实后验。我们引入 PosteriorBench——一个评估生成反问题求解器分布准确性的基准。PosteriorBench 评估四个基于物理的反问题:达西流反演、泊松源恢复、碳捕集与封存以及光传输材质推断。对于每个任务,我们使用计算量大但已确立的程序(如拒绝采样和马尔可夫链蒙特卡洛)构建高保真参考后验,从而直接评估求解器是否恢复完整解集而非单个最优样本。我们将这些参考与五指标后验评估套件配对:后验均值误差、后验标准差误差、最大均值差异、切片 Wasserstein 距离和径向平均功率谱误差。这些指标评估点精度、边缘不确定性、分布对齐和全局频率保真度。基准涵盖稀疏传感、低分辨率观测、非线性正向模型、不同噪声水平和多模态先验,并提供统一的分布匹配和不确定性量化管线。实验揭示了当前求解器中显著的分布匹配差距,同时表明神经算子提高了分辨率鲁棒性,而引导权重和生成噪声是后验方差校准的关键。

原文摘要

Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport m...


*自动采集于 2026-09-19*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens