[论文] MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement...

研究领域: ML 作者: Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch 发布时间: 2026-09-17 arXiv: 2609.20747

论文概要

研究领域: ML 作者: Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch 发布时间: 2026-09-17 arXiv: 2609.20747

中文摘要

强化学习因其超越人类的性能潜力和自学习策略而颇具前景,但它在真实世界自动驾驶中的应用仍然稀少,在非结构化环境中尤其如此,难点在于非结构化环境的 sim-to-real 迁移。本文提出 MILER——一个具备零样本 sim-to-real 迁移能力的端到端策略框架。离线训练阶段,我们使用定制的语义中层表示(MLR)仿真器,用强化学习训练策略网络,控制输出直接作用于自行车模型。真实车辆部署时,相机与激光雷达数据由 BEVFusion 处理,生成与 MLR 仿真器一致的语义鸟瞰图表示;策略网络生成的动作并不直接作用于真实车辆,而是采用轨迹对齐策略,实现感知与控制的零样本迁移。我们在包含多样挑战的测试道路上充分评估:多种障碍物、发卡弯、最高 33.6 km/h 的速度以及越野路段。两辆不同车辆在 3.0 km 测试道上共自动驾驶 17.3 km 无需人工干预,验证了方法的有效性。整套软件栈运行在 Jetson AGX Orin 上。

原文摘要

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation...


*自动采集于 2026-09-20*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens