论文概要
研究领域: ML
作者: Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue
发布时间: 2026-10-02
arXiv: 2610.03702
中文摘要
大语言模型后训练数据的选择对下游性能有重大影响。基于梯度的数据选择是一种流行方法,通过训练数据梯度与小型验证集梯度的对齐程度来排序。然而,使用全参数梯度排序需要对每个样本进行昂贵的前向和反向传播,使得大规模候选池的计算不可行。这引发一个自然的问题:能否以极低成本近似全梯度特征?幸运的是,我们发现输出层梯度足以实现有效的数据选择,且仅需更便宜的前向传播。我们将其实现为 LESSER——一个即插即用的选择方法包装器,将特征提取的 FLOP 成本在 SFT 上降低 9.7 倍,在 RL 基准上降低 3.0 倍,同时在下游任务上保持与全梯度性能一致。实验发现,即使输出层梯度和全梯度对单个样本的排序不同,它们选择的批次仍具有对齐的梯度。
原文摘要
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by \(9.7\times\) for SFT and \(3.0\times\) fo...
自动采集于 2026-10-06
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。