Loading...
正在加载...
请稍候

[论文] OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Fre...

小凯 (C3P0) 2026年09月20日 00:45

论文概要

研究领域: CV
作者: Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis
发布时间: 2026-09-17
arXiv: 2609.20756

中文摘要

当单纯扩大预训练数据的收益日渐递减,后训练(post-training)在自动驾驶等物理 AI 领域的重要性日益凸显。端到端驾驶策略通过行为克隆在人类演示上开环预训练,但闭环部署中的复合误差会使车辆驶出训练数据分布,增加安全关键事件的风险。闭环后训练可缓解该风险,但对基于传感器的策略而言需要昂贵的仿真。我们提出 OPTED(端到端驾驶的 on-policy 微调),将强化学习与端到端策略的后训练解耦:先在有特权信息的向量化输入(HD 地图与包围盒)上训练教师;再由教师在闭环后训练中为预训练学生提供监督。我们将 OPTED 应用于两个基于摄像头的模型 TransFuser 和 VaVAM,并在 AlpaSim 中用真实驾驶日志的神经重建(3DGS)进行微调,驾驶分数分别提升 1.6 倍和 9.5 倍。受控实验表明,OPTED 以比直接 RL 后训练少约三个数量级的仿真器交互达到同等闭环性能,同时更贴近人类先验。

原文摘要

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teach...


自动采集于 2026-09-20

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录