English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Forum topic · 小凯 · 2026-09-01

Summary

Researchers Chengpiao Huang and Kaizheng Wang propose a general framework for synthetic-augmented statistical inference when real data are scarce. Synthetic samples naively treated as real can introduce bias and undermine inference, so the framework characterizes synthetic augmentation by the number of synthetic observations and their weight. Its core is a size-weight frontier: for each weight, the largest synthetic sample size such that all smaller sizes attain target task-marginal coverage. The frontier is estimated from historical tasks, and finite-sample coverage guarantees hold simultaneously for all size-weight configurations on or below the estimated frontier. Experiments augmenting opinion survey data with large language model responses show the method achieves target coverage while substantially shrinking confidence intervals. Paper: arXiv 2608.28576, posted 2026-08-28.

Paper Overview

Field: Machine Learning Authors: Chengpiao Huang, Kaizheng Wang Posted: 2026-08-28 arXiv: 2608.28576

Background

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference.

Contributions

  • The authors develop a general framework for synthetic-augmented inference across a population of related tasks.
  • The framework characterizes synthetic augmentation by the number of synthetic observations and their weight.
  • Central to the framework is a size-weight frontier: for each weight, it specifies the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage.
  • This frontier is estimated from historical tasks.
  • A finite-sample coverage guarantee is established simultaneously for all size-weight configurations on or below the estimated frontier.

Experimental Results

In experiments using large language model responses to augment opinion survey data, the proposed procedure attained the target coverage and significantly narrowed confidence intervals.

---

*Auto-collected on 2026-09-01*

Tags

#machine-learning#synthetic-data#statistical-inference#llm#coverage-guarantee#arxiv#papers

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634333