[论文] QF3: Fast Flow RL with Filtered Q-Gradients

研究领域: ML 作者: Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Chung Min Kim, Brent Yi, David McAllister, Hongsuk Choi, Himanshu Gaurav Singh, Jinkun Cao, Ken Goldberg, Pieter Abbeel, Carmelo Sferrazza, Angjoo Kanazawa 发布时间: 2026-10-06 arXiv: 2610.08789

中文摘要

Flow策略已成为从示范中学习机器人行为的标准策略类别,但强化学习对于改进预训练的flow策略或通过交互从头学习它们仍然至关重要。我们引入QF3(带过滤Q梯度的快速Flow强化学习),这是一种在线离线策略RL算法,通过flow matching加上critic的动作梯度来训练flow策略,该梯度通过flow输出的一步预测进行反向传播。为了保持critic和该预测的可靠性,QF3仅将critic梯度应用于保持在回放缓冲动作附近的动作维度。据我们所知,QF3是第一个能够从头训练人形运动策略并将其零样本迁移到硬件的离线策略flow RL方法。配合高吞吐量离线策略训练方案,它以比最近的在线策略flow RL方法FPO++快10倍的训练速度来训练人形运动和动作跟踪策略。我们还进一步将QF3应用于在ABC-Sim和Robomimic任务上微调预训练的基于flow的操作策略。这些结果表明QF3既能从头学习机器人策略,也能改进从示范中获取的策略。

原文摘要

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with ...


*自动采集于 2026-10-08*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens