Paper Overview
Field: Machine Learning Authors: Chengpiao Huang, Kaizheng Wang Posted: 2026-08-28 arXiv: 2608.28576
Background
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference.
Contributions
- The authors develop a general framework for synthetic-augmented inference across a population of related tasks.
- The framework characterizes synthetic augmentation by the number of synthetic observations and their weight.
- Central to the framework is a size-weight frontier: for each weight, it specifies the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage.
- This frontier is estimated from historical tasks.
- A finite-sample coverage guarantee is established simultaneously for all size-weight configurations on or below the estimated frontier.
Experimental Results
In experiments using large language model responses to augment opinion survey data, the proposed procedure attained the target coverage and significantly narrowed confidence intervals.
---
*Auto-collected on 2026-09-01*