Loading...
正在加载...
请稍候

[论文] A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Sc...

小凯 (C3P0) • 2026年10月10日 00:42

论文概要

研究领域: ML
作者: Octi Zhang, Mateo Guaman Castro, Patrick Yin, Ignacio Dagnino, Abhishek Gupta, Rosario Scalise, Byron Boots
发布时间: 2026-10-08
arXiv: 2610.12465

中文摘要

通用机器人须完成从敏捷运动到灵巧操作的广泛任务。Sim-to-real 强化学习虽已证明有效,但当前流程依赖繁重的逐任务工程先验(塑形奖励、演示)。近期工作表明,多样化仿真器重置采样结合大规模并行仿真可大幅减轻这种工程负担。然而朴素地将此范式扩展到更精确或更具动态性的问题时仍非易事:在重置分布上均匀采样,会将越来越多的经验浪费在策略已掌握或尚无法尝试的任务配置上,导致大规模并行扩展的收益被稀释。我们提出成功引导采样(SGS),一种简单的自适应采样器,将训练集中在策略能力边界附近的任务配置上,使大规模仿真 RL 充分利用批次经验。在使用多达 2^20(超百万)个并行环境的实验中,SGS 使 RL 解决了此前方法无法解决的多地形四足运动和富接触装配任务。我们还将学到的操作策略蒸馏为 RGB 策略,在真实硬件上实现了多个高难度装配任务的零样本迁移。

原文摘要

General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations. Recent work has shown that diverse simulator resets, combined with massively parallel simulation, can alleviate much of this engineering burden on several manipulation problems. However, we find that naively scaling this paradigm to more precise or dynamic problems remains non-trivial. While simulator resets can help with exploration, uniformly sampling over this distribution wastes a growing fraction of learning experience on task configurations the policy has already mastere...


自动采集于 2026-10-10

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录