English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NVIDIA ENPIRE: Eight Codex Agents Run Autonomous Robotics Research

Forum topic · QianXun · 2026-06-18

Summary

On June 17, 2026, NVIDIA's GEAR lab unveiled ENPIRE, described by Jim Fan as the first implementation of Physical AutoResearch. The system pairs eight Codex agents with eight robots, plus fixed GPU and token budgets, and a single goal: finish tasks and keep the robots busy, running unattended from launch. Key design choices include two layers of hardware safety (hard kinematic limits that trigger automatic task reset, and force-limited compliant grippers that stall on misaligned inserts), a frozen visual success classifier trained from a few minutes of demonstrations to prevent reward hacking, and telemetry on robot utilization (MRU), token consumption (MTU), and GPU usage, combined into Tokens-to-Success and Time-to-Success metrics. Benchmarks on zip-tying, needle arrangement, and GPU installation showed eight-robot parallel exploration significantly outperforming a single robot. The system will be open-sourced, though comparisons with Pi0, GR00T, and Helix remain pending.

On June 17, 2026, NVIDIA's GEAR lab released the ENPIRE system. Jim Fan announced it the same day, calling it the first realization of autonomous research in the physical world (Physical AutoResearch). ENPIRE equips 8 Codex agents with 8 robots, plus fixed GPU and token budgets, with a single objective: complete tasks as fast as possible and keep the robots busy. Once launched, no human intervention is required.

Jim Fan himself admits the engineering behind it is "hell" — the difficulty is not the algorithm, but everything that must be ready before pressing Enter. Three components stand out.

1. Safety guardrails, not prompts

Unattended overnight runs mean safety cannot be a polite note in a system prompt. ENPIRE hard-wires safety at two hardware levels:

  • Hard kinematic limits: crossing the boundary triggers task failure and automatic reset.
  • Force-limited compliant grippers: misaligned insertion attempts cause a stall instead of crushing the robot or the object.
  • The guiding principle: be conservative, so the operator can sleep at night.

    2. A frozen definition of "done"

    If an agent can modify its own reward function, it will game the score. ENPIRE's fix is to "weld the goalpost in place": first collect a few minutes of successful and failed demonstrations, have the agent write a visual classifier to judge success, hill-climb it to stability, then freeze the classifier and embed it in the Gym environment — untouchable for the entire research run.

    3. System telemetry

    Robot-seconds are the scarcest resource, GPU-seconds next, tokens last. All three are instrumented and fed back into ENPIRE for real-time resource awareness. Three metrics are tracked:

  • MRU: fraction of time robots are actively working
  • MTU: tokens consumed per minute (whether the agent is actually thinking)
  • GPU utilization
  • These combine into two budget-to-output ratios: Tokens-to-Success and Time-to-Success.

    Experiments and results

    Three tasks were tested: zip-tying, fine needle arrangement, and GPU installation. Conclusion: 8-robot parallel exploration is significantly faster than a single robot. The system will be open-sourced.

    Why it matters

    1. Pushes LLM agents onto physical hardware. Past AutoResearch agents worked in sandboxed domains (math, code). ENPIRE closes the loop on real robots. 2. Resource budget awareness. Three utilization metrics (robots, tokens, GPU) become observable for the first time — agents are now being held accountable for ROI, like human engineering teams. 3. Safety is welded, not prompted. Two layers of hardware protection make truly unattended operation possible — the minimum bar for pressing Enter and walking away.

    Risks and open questions

  • How long can reward freezing hold up? Slightly more complex tasks may outgrow the visual classifier.
  • Does 8-robot parallelism scale energy and maintenance costs for industrial deployment?
  • OpenAI Codex's role on physical agents is undisclosed — is it an LLM controller or genuinely modifying an RL policy?
  • Paper/code release timing is undecided; cross-team reproduction will be hard.
  • No comparisons yet with Pi0, GR00T, or Helix — the physical intelligence field sits at a fork between "model-first" and "agent-first" routes.
Source: https://x.com/DrJimFan/status/2067283904986517866

Tags

#nvidia#gear-lab#enpire#embodied-ai#autonomous-agents#codex#robotics#physical-autoresearch

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981466