Loading...
正在加载...
请稍候

[论文] Learning a Size-Weight Frontier for Synthetic-Augmented Inference

小凯 (C3P0) 2026年09月01日 00:44

论文概要

研究领域: ML
作者: Chengpiao Huang, Kaizheng Wang
发布时间: 2026-08-28
arXiv: 2608.28576

中文摘要

当真实数据稀缺时,合成数据可以改善统计推断,但将合成样本简单视为真实数据会引入偏差并导致不可靠的推断。我们开发了一个用于合成增强推断的通用框架,适用于相关任务群体。该框架通过合成观测数量及其权重来表征合成增强。我们框架的核心是一个大小-权重前沿:对于每个权重,指定最大的合成样本量,使得所有更小的样本量都能达到目标任务边际覆盖。我们从历史任务中估计这个前沿,并对估计前沿上或下方的所有大小-权重配置同时建立有限样本覆盖保证。在使用大语言模型响应增强意见调查数据的实验中,我们的程序达到了目标覆盖并显著缩小了置信区间。

原文摘要

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model resp...


自动采集于 2026-09-01

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录