论文概要
研究领域: CV
作者: Damiano Da Col, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, Christos Sakaridis
发布时间: 2026-09-17
arXiv: 2609.20756
中文摘要
当单纯扩大预训练数据的收益日渐递减,后训练(post-training)在自动驾驶等物理 AI 领域的重要性日益凸显。端到端驾驶策略通过行为克隆在人类演示上开环预训练,但闭环部署中的复合误差会使车辆驶出训练数据分布,增加安全关键事件的风险。闭环后训练可缓解该风险,但对基于传感器的策略而言需要昂贵的仿真。我们提出 OPTED(端到端驾驶的 on-policy 微调),将强化学习与端到端策略的后训练解耦:先在有特权信息的向量化输入(HD 地图与包围盒)上训练教师;再由教师在闭环后训练中为预训练学生提供监督。我们将 OPTED 应用于两个基于摄像头的模型 TransFuser 和 VaVAM,并在 AlpaSim 中用真实驾驶日志的神经重建(3DGS)进行微调,驾驶分数分别提升 1.6 倍和 9.5 倍。受控实验表明,OPTED 以比直接 RL 后训练少约三个数量级的仿真器交互达到同等闭环性能,同时更贴近人类先验。
原文摘要
As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teach...
自动采集于 2026-09-20
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。