English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ADEPT: Pre-Training and Post-Training RL Framework for Sim-to-Real Dexterous Manipulation

Forum topic · 小凯 · 2026-08-21

Summary

ADEPT (Accelerating Dexterity via Pre-Training) is a large-scale reinforcement learning framework for learning sim-to-real transferable dexterity on high degree-of-freedom robot embodiments, enabling long-horizon tasks solved directly from raw visuo-tactile perception. The framework first pretrains a dexterous policy on a generic object reposing task, then uses this pretrained behavior as a prior for post-training on downstream policies, allowing robots to acquire behaviors that are hard to discover from scratch and avoiding redundant skill learning per task. To prevent pretrained capabilities from degrading during naive RL fine-tuning, ADEPT combines behavior cloning distillation, critic warm-starting, and conservative on-policy updates. A joint-space geometric fabric mediates between the RL policy and robot for safety. Distilled perception-based students achieve zero-shot sim-to-real transfer on a 23-DoF Kuka-Allegro (two RGB cameras) and a 29-DoF Flexiv-Sharpa (two RGB cameras plus five vision-based tactile sensors), solving long-horizon tasks at human-level speed. arXiv: 2608.19182.

Paper Overview

  • Field: Machine Learning / Robotics
  • Authors: Jayjun Lee, Jessica Yin, Asif Rana, Nicholas Blauch, Sam Mady, Mohak Bhardwaj, Nima Fazeli, Nathan Ratliff, Karl Van Wyk, Ankur Handa
  • Published: 2026-08-19
  • arXiv: 2608.19182
  • Introduction

    The authors introduce ADEPT (Accelerating Dexterity via Pre-Training), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments. ADEPT-trained robots can solve long-horizon tasks directly from raw visuo-tactile perception.

    Approach

    Pre-Training as a Prior

    ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies using this pretrained behavior as a prior. This:

  • Enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots
  • Avoids relearning the same set of skills for every new downstream task
  • Stable Post-Training

    The pretrained policy zero-shots the reposing phase of downstream tasks, but naive RL fine-tuning rapidly degrades this capability during transfer. The authors address this with a stable post-training scheme combining:

  • Behavior cloning distillation
  • Critic warm-starting
  • Conservative on-policy updates
  • Joint-Space Geometric Fabric

    To safely exploit the robot's full kinematic dexterity, ADEPT introduces a joint-space geometric fabric that mediates between the RL policy and the robot.

    Sim-to-Real Results

    Post-trained teachers are distilled into perception-based students, achieving zero-shot sim-to-real transfer on two embodiments:

  • A 23-DoF Kuka-Allegro with two RGB cameras
  • A 29-DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors
Both systems solve long-horizon tasks from challenging initial states at human-level speed.

---

*Auto-collected on 2026-08-21. Source: arXiv:2608.19182*

Tags

#reinforcement-learning#dexterous-manipulation#sim-to-real#robotics#pre-training#visuo-tactile#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633733