[论文] Learning a Size-Weight Frontier for Synthetic-Augmented Inference
研究领域: ML 作者: Chengpiao Huang, Kaizheng Wang 发布时间: 2026-08-28 arXiv: 2608.28576
论文概要
研究领域: ML 作者: Chengpiao Huang, Kaizheng Wang 发布时间: 2026-08-28 arXiv: 2608.28576
中文摘要
当真实数据稀缺时,合成数据可以改善统计推断,但将合成样本简单视为真实数据会引入偏差并导致不可靠的推断。我们开发了一个用于合成增强推断的通用框架,适用于相关任务群体。该框架通过合成观测数量及其权重来表征合成增强。我们框架的核心是一个大小-权重前沿:对于每个权重,指定最大的合成样本量,使得所有更小的样本量都能达到目标任务边际覆盖。我们从历史任务中估计这个前沿,并对估计前沿上或下方的所有大小-权重配置同时建立有限样本覆盖保证。在使用大语言模型响应增强意见调查数据的实验中,我们的程序达到了目标覆盖并显著缩小了置信区间。
原文摘要
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model resp...
*自动采集于 2026-09-01*
#论文 #arXiv #ML #小凯