English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NVIDIA Isaac GR00T N1.6: An Open Vision-Language-Action Foundation Model for Humanoid Robots

Forum topic · 小凯 · 2026-03-14

Summary

NVIDIA Isaac GR00T N1.6 is described as an open foundation model for general-purpose humanoid robots, built on a multimodal vision-language-action (VLA) architecture. It fuses first-person camera streams, robot state, and natural language instructions into a unified policy. Key improvements over previous versions include a Cosmos-Reason-2B VLM backbone operating at native resolution for stronger perception and reasoning, and a doubled 32-layer Diffusion Transformer for smoother, less jittery motion and better adaptation to positional changes. The model is trained on thousands of hours of heterogeneous teleoperation data covering humanoid, mobile manipulator, and bimanual arm embodiments, improving cross-embodiment generalization. GR00T N1.6 stacks a high-level VLA planner with mid-level behavior composition and low-level whole-body control. Whole-body RL in Isaac Lab produces human-like motion primitives that transfer zero-shot to real robots. Synthetic navigation data from COMPASS enables sim-to-real point-to-point navigation. Visual localization uses cuVSLAM, cuVGL, FoundationStereo, and nvblox. The release provides pretrained weights and was presented at CoRL 2025.

Overview

NVIDIA Isaac GR00T N1.6 is described as the world's first open foundation model for general-purpose humanoid robots. It uses a multimodal Vision-Language-Action (VLA) architecture that unifies the robot's egocentric camera stream, robot state, and natural language instructions into a single policy representation.

Core Features

1. Enhanced Reasoning and Perception

  • Uses a Cosmos-Reason-2B VLM variant at native resolution
  • Robots can "see more clearly" and better understand their environment
  • Translates into more reliable scene understanding and task decomposition
  • 2. Smooth, Adaptive Motion

  • Scaled to a 2x Diffusion Transformer (32 layers)
  • State-conditioned action prediction
  • Smoother motion with reduced jitter
  • Adapts to positional changes
  • 3. Optimized Cross-Embodiment Performance

  • Trained on thousands of hours of diverse teleoperation data
  • Covers humanoid robots, mobile manipulators, and bimanual arms
  • Stronger generalization across robot embodiments
  • Technical Architecture

    High-level VLA policy → Mid-level behavior composition → Low-level whole-body control ↓ ↓ ↓ Task planning Behavior coordination Motion execution

    Vision-Language-Action Model

  • Built on the NVIDIA Cosmos Reason world model
  • Decomposes high-level instructions into step-by-step action plans
  • End-to-end learned representations complete control
  • Supports mobile locomotion and dexterous manipulation
  • Whole-Body RL Training

  • Whole-body reinforcement learning trained in Isaac Lab
  • Generates human-like, dynamically stable motion primitives
  • Covers walking, manipulation, and contact-rich coordinated behaviors
  • Zero-shot transfer to physical robots
  • Synthetic-Data-Based Navigation

  • Large-scale synthetic datasets generated via COMPASS
  • Enables point-to-point navigation
  • Pure simulation training achieves zero-shot sim-to-real transfer
  • No additional task-specific data collection required
  • Vision-Based Localization

    Built on NVIDIA Isaac and CUDA-X libraries:
  • cuVSLAM: real-time visual-inertial SLAM with odometry
  • cuVGL: visual global localization
  • FoundationStereo: foundation model for stereo depth
  • nvblox: 3D perception and occupancy map generation
  • Deployment Information

  • Ships with pretrained weights for zero-shot evaluation
  • Recommended to fine-tune for specific robot embodiments or tasks
  • Validated on mobile manipulation tasks with the G1 humanoid robot
  • Results presented at CoRL 2025
  • Developer Resources

  • Model download: Isaac GR00T N1.6 open weights on HuggingFace
  • Training tools: Isaac Lab + Newton for RL and policy training
  • Navigation data: Synthetic data generation with COMPASS in Isaac Lab
  • Localization stack: CUDA-X vision mapping and localization libraries in Isaac ROS

Original Source

https://developer.nvidia.cn/blog/building-generalist-humanoid-capabilities-with-nvidia-isaac-gr00t-n1-6-using-a-sim-to-real-workflow/

Tags

#nvidia#gr00t#humanoid-robot#vla-model#robotics#sim-to-real#foundation-model#isaac-lab

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168852