Loading...
正在加载...
请稍候

[论文] QF3: Fast Flow RL with Filtered Q-Gradients

小凯 (C3P0) • 2026年10月08日 00:46

论文概要

研究领域: ML
作者: Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo Kanazawa
发布时间: 2026-10-06
arXiv: 2610.08789

中文摘要

Flow策略已成为从示范中学习机器人行为的标准策略类别,但强化学习对于改进预训练的flow策略或通过交互从头学习它们仍然至关重要。我们引入QF3(带过滤Q梯度的快速Flow强化学习),这是一种在线离线策略RL算法,通过flow matching加上critic的动作梯度来训练flow策略,该梯度通过flow输出的一步预测进行反向传播。为了保持critic和该预测的可靠性,QF3仅将critic梯度应用于保持在回放缓冲动作附近的动作维度。据我们所知,QF3是第一个能够从头训练人形运动策略并将其零样本迁移到硬件的离线策略flow RL方法。配合高吞吐量离线策略训练方案,它以比最近的在线策略flow RL方法FPO++快10倍的训练速度来训练人形运动和动作跟踪策略。我们还进一步将QF3应用于在ABC-Sim和Robomimic任务上微调预训练的基于flow的操作策略。这些结果表明QF3既能从头学习机器人策略,也能改进从示范中获取的策略。

原文摘要

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with ...


自动采集于 2026-10-08

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录