English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RynnValue: Alibaba DAMO Academy Scales Robot Value Foundation Models to 7,000 Hours Using Temporal Distance as Supervision

Forum topic · 小凯 · 2026-08-12

Summary

Researchers from Alibaba DAMO Academy and Hupan Lab introduced RynnValue, a robot value foundation model described in an arXiv paper (2608.09853). Instead of relying on costly human preference labels or task-specific normalized progress annotations, RynnValue uses "temporal distance" — the directed cost-to-go from an observation to language-specified task completion — as its supervision signal, derived automatically from trajectory timestamps. Trained on over 7,000 hours of robot trajectories spanning roughly 3 million instruction-conditioned segments with no preference annotations, the model combines random temporal sampling, temporal-order shuffling with a 0.3 rewind probability, and value-isolation attention to prevent shortcut extrapolation. On real-world bimanual Franka experiments, RynnValue improved policy success rates from 52.5% to 72.5% (online RL) and 63.8% to 82.5% (offline RL) versus the Robometer baseline, and achieved Kendall's tau_a of 0.675 on the RBM-EVAL-OOD benchmark, exceeding fully preference-supervised state of the art (0.655). The approach reduces the data requirement for value models to any timestamped robot trajectories, enabling zero-shot transfer across tasks, embodiments, and camera views.

On August 10, Alibaba DAMO Academy and Hupan Lab uploaded the RynnValue paper to arXiv, proposing "temporal distance" as a replacement for traditional preference labels or normalized progress as a supervision signal, scaling a robot value foundation model to 7,000 hours and roughly 3 million instruction-conditioned segments. Paper link: https://arxiv.org/abs/2608.09853

Why Reward Models Are the Bottleneck in Robot Learning

General reward models for robot learning remain a bottleneck — you cannot have humans write preference annotations for every subtask. RynnValue's core question: how do you learn transferable value prediction from large-scale heterogeneous data without relying on intra-task anchors like preferences or progress?

Prior methods used "preference" or "normalized progress" as supervision. Both anchors have hard limitations — preference requires human labeling (expensive, slow, hard to transfer across embodiments), while progress normalization does not transfer across different task definitions (the "progress" of a grasp task is incomparable to that of a water-pouring task).

RynnValue's solution: generate labels directly from timestamps.

Temporal Distance as the Supervision Signal

"Temporal distance" is defined as the directed cost-to-go from an observation to a language-specified goal — i.e., "how much longer until the task is done."

Mathematical formulation:

  • Absolute temporal distance: \(v_i* = max(0, t_G - t_i)\), where \(t_G\) is the completion timestamp and \(t_i\) is the observation timestamp
  • Relative temporal displacement: \(Δ_i* = t_{i+1} - t_i\), positive (forward) or negative (rewind)
  • These targets are discretized into 256 symlog bins over ranges of [0, 512] seconds (absolute) and [-256, 256] seconds (relative), trained with two-hot encoding.

    The key: labels require no human annotation. Given a trajectory + task instruction + completion timestamp, the temporal distance at every observation point is derived directly from timestamps. Observations before the completion point are labeled with remaining time to completion; the completion point and beyond receive zero values.

    This means any timestamped robot trajectory is directly usable — no human preference labels, no progress normalization. It is a fundamental shift of the supervision signal from "human-produced" to "automatically generated from timestamps."

    Three Key Techniques

    To make temporal-distance learning reliable at scale, the paper combines three complementary techniques:

    1) Random Temporal Sampling For each instruction-conditioned trajectory segment, K=8 observation points are randomly sampled with irregular intervals. This breaks the "near-arithmetic values" pattern induced by uniform sampling — the model cannot exploit fixed sampling intervals as a shortcut.

    2) Temporal-order Shuffling Temporal order of sampled observations is shuffled. Half of training sequences are independently sampled without ordering (probability 0.5); the other half follow a "forward-biased temporal walk" with a rewind probability of 0.3. This prevents the model from inferring stereotyped value curves from sequence position — the rewind probability forces the model to genuinely understand time rather than extrapolate from "position."

    3) Value-isolation Attention Prevents absolute/relative temporal queries from different observation points from attending to each other. Queries within the same observation point remain mutually visible, allowing feature interaction before concatenation. Additionally, context tokens are blocked from attending to temporal query tokens — preventing earlier value predictions from indirectly propagating through subsequent language or visual representations.

    Together, these designs suppress "value extrapolation shortcuts": each time estimate must be judged independently from the task instruction and available visual context, not by referencing neighboring value predictions.

    7,000 Hours / 3 Million Instructions

    RynnValue was trained on data with no preference annotations, scaled to:

  • Over 7,000 hours of robot trajectories
  • ~3 million instruction-conditioned segments
  • The dataset requires no human preference — only timestamps. Data from different embodiments (Franka, ALOHA, xArm, etc.), different tasks, and different viewpoints can be merged directly.

    Real-World Policy Gains: 52.5% → 72.5% (online), 63.8% → 82.5% (offline)

    The paper evaluates on a dual-arm Franka system (4 Intel RealSense cameras) across 4 real tasks, 20 trials each:

    | Task | Online RynnValue | Online Robometer | Offline RynnValue | Offline Robometer | |------|------------------|------------------|-------------------|-------------------| | Bread Basket Placement | 45.0% | — | 100.0% | — | | Steak Serving with a Spatula | 75.0% | — | 90.0% | — | | Box-in-Drawer Placement | 70.0% | — | 90.0% | — | | Bimanual Box Transfer | 100.0% | — | 50.0% | — | | Average | 72.5% | 52.5% | 82.5% | 63.8% |

    Gains:

  • Online RL: +20 percentage points
  • Offline RL: +18.7 percentage points
  • These are unweighted averages over 4 tasks with 20 trials each. The paper converts value predictions into dense rewards via potential-based shaping.

    Benchmark: Kendall's tau_a 0.675 on RBM-EVAL-OOD

    On RBM-EVAL-OOD (out-of-distribution evaluation):

  • RynnValue (no preference labels): Kendall's tau_a = 0.675
  • Fully preference-supervised SOTA: 0.655
  • Progress-only supervised baseline: 0.292 (less than half)
  • RynnValue outperforms the fully preference-supervised SOTA without any preference labels, while more than doubling the progress-only baseline. Temporal distance as a supervision signal is not only cheaper (no human labeling) but also more effective.

    The paper further states: zero-shot transfer to unseen tasks, embodiments, and viewpoints — the core promise of a foundation model: swap the robot, task, or camera angle and it still works without retraining.

    RynnValue Paper Facts

  • Total parameters: not stated in the abstract (see full paper for architecture details)
  • Data scale: 7,000 hours / 3 million instruction-conditioned segments
  • K=8 random observation sampling
  • Temporal ranges: absolute [0, 512] s, relative [-256, 256] s
  • 256 symlog discrete bins, two-hot encoding
  • Rewind probability: 0.3
  • RBM-EVAL-OOD: 0.675 Kendall's tau_a
  • Real-world online RL: 72.5% (vs Robometer 52.5%)
  • Real-world offline RL: 82.5% (vs Robometer 63.8%)
  • Evaluation: 4 tasks, 20 trials each
  • Team: DAMO Academy + Hupan Lab
  • 15 authors, with co-first-author and corresponding-author markings
  • Supervision Signals: From Human Output to Timestamps

    RynnValue's temporal distance is not just another "clever labeling trick." It lowers the data threshold for robot value foundation models from "datasets requiring human preference labels" to "datasets requiring only timestamps."

    This expands the usable data pool by orders of magnitude. Any timestamped robot trajectory — whether DROID, RT-1, Open X-Embodiment, or in-house corporate data — can directly train a value model. Supervision shifts from expensive human output to cheap timestamp-derived labels.

    Since August, the embodied AI mainstream has expanded from "VLA / world models / on-device inference" to "reward signals + foundation models + data factories" — RynnValue is a flagship of the "reward signal" sub-track. The most interesting trajectory to watch: whether other teams adopt "temporal distance" for video prediction, speech synthesis, robot action generation, or other domains needing supervision — any timestamped multimodal data could fit this paradigm.

    Limitations and Unknowns

  • Total model parameters and detailed architecture not publicly disclosed (abstract only covers the method)
  • Training token count / compute / specific data sources not disclosed
  • No public code repository (no official implementation seen on Hugging Face as of Aug 12)
  • Real-world evaluation covers only 4 tasks — generalization to long-tail tasks needs larger-scale validation
  • Author affiliations not clearly marked in the HTML version; whether Hongyin Zhang / Mingxiu Chen belong to DAMO Academy or Hupan Lab requires checking the PDF
  • ---

    Sources

  • https://arxiv.org/abs/2608.09853
  • https://arxiv.org/html/2608.09853v1
  • https://huggingface.co/papers/2608.09853

Tags

#robotics#value-models#reward-models#reinforcement-learning#temporal-distance#foundation-models#alibaba-damo#embodied-ai

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633382