Summary
PosteriorBench (arXiv:2609.20794) is a benchmark for evaluating the distributional accuracy of generative models used to solve scientific inverse problems. Existing evaluations focus on whether a method produces a single plausible reconstruction, which is insufficient for ill-posed problems where multiple solutions may be consistent with the same sparse or noisy observations. A solver can achieve strong pointwise accuracy while failing to capture the true posterior via mode collapse, overconfident uncertainty, or averaging incompatible solutions. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, high-fidelity reference posteriors are constructed using established but computationally expensive methods such as rejection sampling and Markov chain Monte Carlo. The benchmark pairs these references with a five-metric posterior evaluation suite: posterior mean error, posterior standard deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power spectrum error. Experiments reveal significant distribution-matching gaps in current solvers, while showing that neural operators improve resolution robustness, and that guidance weights and generation noise are key to posterior variance calibration.
Overview
Research area: Machine Learning
Authors: Jiachen Yao, Zi-Siang Hsu, Xi Deng, Aditi Gupta, Xin Ju, Sally M Benson, Gege Wen, Anima Anandkumar
arXiv: 2609.20794
Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions.
Key contributions
- PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers, covering four physics-based inverse problems:
- Darcy flow inversion
- Poisson source recovery
- Carbon capture and storage
- Light transport material inference
- For each task, high-fidelity reference posteriors are constructed using computationally expensive but established procedures (e.g., rejection sampling and Markov chain Monte Carlo), enabling direct evaluation of whether a solver recovers the full solution set rather than a single best sample.
- A five-metric posterior evaluation suite:
- Posterior mean error
- Posterior standard deviation error
- Maximum mean discrepancy (MMD)
- Sliced Wasserstein distance
- Radially averaged power spectrum error
These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark covers sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, and provides a unified pipeline for distribution matching and uncertainty quantification.
Findings
Experiments reveal significant distribution-matching gaps in current solvers. Neural operators improve robustness to resolution changes, while guidance weights and generation noise are critical factors for posterior variance calibration.
*Source: zhichai.net, auto-collected 2026-09-19.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178634985