[论文] Learning a Size-Weight Frontier for Synthetic-Augmented Inference

研究领域: ML 作者: Chengpiao Huang, Kaizheng Wang 发布时间: 2026-08-28 arXiv: 2608.28576

论文概要

研究领域: ML 作者: Chengpiao Huang, Kaizheng Wang 发布时间: 2026-08-28 arXiv: 2608.28576

中文摘要

当真实数据稀缺时,合成数据可以改善统计推断,但将合成样本简单视为真实数据会引入偏差并导致不可靠的推断。我们开发了一个用于合成增强推断的通用框架,适用于相关任务群体。该框架通过合成观测数量及其权重来表征合成增强。我们框架的核心是一个大小-权重前沿:对于每个权重,指定最大的合成样本量,使得所有更小的样本量都能达到目标任务边际覆盖。我们从历史任务中估计这个前沿,并对估计前沿上或下方的所有大小-权重配置同时建立有限样本覆盖保证。在使用大语言模型响应增强意见调查数据的实验中,我们的程序达到了目标覆盖并显著缩小了置信区间。

原文摘要

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model resp...


*自动采集于 2026-09-01*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens