On August 10, Alibaba DAMO Academy and Hupan Lab uploaded the RynnValue paper to arXiv, proposing "temporal distance" as a replacement for traditional preference labels or normalized progress as a supervision signal, scaling a robot value foundation model to 7,000 hours and roughly 3 million instruction-conditioned segments. Paper link: https://arxiv.org/abs/2608.09853
Why Reward Models Are the Bottleneck in Robot Learning
General reward models for robot learning remain a bottleneck — you cannot have humans write preference annotations for every subtask. RynnValue's core question: how do you learn transferable value prediction from large-scale heterogeneous data without relying on intra-task anchors like preferences or progress?
Prior methods used "preference" or "normalized progress" as supervision. Both anchors have hard limitations — preference requires human labeling (expensive, slow, hard to transfer across embodiments), while progress normalization does not transfer across different task definitions (the "progress" of a grasp task is incomparable to that of a water-pouring task).
RynnValue's solution: generate labels directly from timestamps.
Temporal Distance as the Supervision Signal
"Temporal distance" is defined as the directed cost-to-go from an observation to a language-specified goal — i.e., "how much longer until the task is done."
Mathematical formulation:
- Absolute temporal distance: \(v_i* = max(0, t_G - t_i)\), where \(t_G\) is the completion timestamp and \(t_i\) is the observation timestamp
- Relative temporal displacement: \(Δ_i* = t_{i+1} - t_i\), positive (forward) or negative (rewind)
- Over 7,000 hours of robot trajectories
- ~3 million instruction-conditioned segments
- Online RL: +20 percentage points
- Offline RL: +18.7 percentage points
- RynnValue (no preference labels): Kendall's tau_a = 0.675
- Fully preference-supervised SOTA: 0.655
- Progress-only supervised baseline: 0.292 (less than half)
- Total parameters: not stated in the abstract (see full paper for architecture details)
- Data scale: 7,000 hours / 3 million instruction-conditioned segments
- K=8 random observation sampling
- Temporal ranges: absolute [0, 512] s, relative [-256, 256] s
- 256 symlog discrete bins, two-hot encoding
- Rewind probability: 0.3
- RBM-EVAL-OOD: 0.675 Kendall's tau_a
- Real-world online RL: 72.5% (vs Robometer 52.5%)
- Real-world offline RL: 82.5% (vs Robometer 63.8%)
- Evaluation: 4 tasks, 20 trials each
- Team: DAMO Academy + Hupan Lab
- 15 authors, with co-first-author and corresponding-author markings
- Total model parameters and detailed architecture not publicly disclosed (abstract only covers the method)
- Training token count / compute / specific data sources not disclosed
- No public code repository (no official implementation seen on Hugging Face as of Aug 12)
- Real-world evaluation covers only 4 tasks — generalization to long-tail tasks needs larger-scale validation
- Author affiliations not clearly marked in the HTML version; whether Hongyin Zhang / Mingxiu Chen belong to DAMO Academy or Hupan Lab requires checking the PDF
- https://arxiv.org/abs/2608.09853
- https://arxiv.org/html/2608.09853v1
- https://huggingface.co/papers/2608.09853
These targets are discretized into 256 symlog bins over ranges of [0, 512] seconds (absolute) and [-256, 256] seconds (relative), trained with two-hot encoding.
The key: labels require no human annotation. Given a trajectory + task instruction + completion timestamp, the temporal distance at every observation point is derived directly from timestamps. Observations before the completion point are labeled with remaining time to completion; the completion point and beyond receive zero values.
This means any timestamped robot trajectory is directly usable — no human preference labels, no progress normalization. It is a fundamental shift of the supervision signal from "human-produced" to "automatically generated from timestamps."
Three Key Techniques
To make temporal-distance learning reliable at scale, the paper combines three complementary techniques:
1) Random Temporal Sampling For each instruction-conditioned trajectory segment, K=8 observation points are randomly sampled with irregular intervals. This breaks the "near-arithmetic values" pattern induced by uniform sampling — the model cannot exploit fixed sampling intervals as a shortcut.
2) Temporal-order Shuffling Temporal order of sampled observations is shuffled. Half of training sequences are independently sampled without ordering (probability 0.5); the other half follow a "forward-biased temporal walk" with a rewind probability of 0.3. This prevents the model from inferring stereotyped value curves from sequence position — the rewind probability forces the model to genuinely understand time rather than extrapolate from "position."
3) Value-isolation Attention Prevents absolute/relative temporal queries from different observation points from attending to each other. Queries within the same observation point remain mutually visible, allowing feature interaction before concatenation. Additionally, context tokens are blocked from attending to temporal query tokens — preventing earlier value predictions from indirectly propagating through subsequent language or visual representations.
Together, these designs suppress "value extrapolation shortcuts": each time estimate must be judged independently from the task instruction and available visual context, not by referencing neighboring value predictions.
7,000 Hours / 3 Million Instructions
RynnValue was trained on data with no preference annotations, scaled to:
The dataset requires no human preference — only timestamps. Data from different embodiments (Franka, ALOHA, xArm, etc.), different tasks, and different viewpoints can be merged directly.
Real-World Policy Gains: 52.5% → 72.5% (online), 63.8% → 82.5% (offline)
The paper evaluates on a dual-arm Franka system (4 Intel RealSense cameras) across 4 real tasks, 20 trials each:
| Task | Online RynnValue | Online Robometer | Offline RynnValue | Offline Robometer | |------|------------------|------------------|-------------------|-------------------| | Bread Basket Placement | 45.0% | — | 100.0% | — | | Steak Serving with a Spatula | 75.0% | — | 90.0% | — | | Box-in-Drawer Placement | 70.0% | — | 90.0% | — | | Bimanual Box Transfer | 100.0% | — | 50.0% | — | | Average | 72.5% | 52.5% | 82.5% | 63.8% |
Gains:
These are unweighted averages over 4 tasks with 20 trials each. The paper converts value predictions into dense rewards via potential-based shaping.
Benchmark: Kendall's tau_a 0.675 on RBM-EVAL-OOD
On RBM-EVAL-OOD (out-of-distribution evaluation):
RynnValue outperforms the fully preference-supervised SOTA without any preference labels, while more than doubling the progress-only baseline. Temporal distance as a supervision signal is not only cheaper (no human labeling) but also more effective.
The paper further states: zero-shot transfer to unseen tasks, embodiments, and viewpoints — the core promise of a foundation model: swap the robot, task, or camera angle and it still works without retraining.
RynnValue Paper Facts
Supervision Signals: From Human Output to Timestamps
RynnValue's temporal distance is not just another "clever labeling trick." It lowers the data threshold for robot value foundation models from "datasets requiring human preference labels" to "datasets requiring only timestamps."
This expands the usable data pool by orders of magnitude. Any timestamped robot trajectory — whether DROID, RT-1, Open X-Embodiment, or in-house corporate data — can directly train a value model. Supervision shifts from expensive human output to cheap timestamp-derived labels.
Since August, the embodied AI mainstream has expanded from "VLA / world models / on-device inference" to "reward signals + foundation models + data factories" — RynnValue is a flagship of the "reward signal" sub-track. The most interesting trajectory to watch: whether other teams adopt "temporal distance" for video prediction, speech synthesis, robot action generation, or other domains needing supervision — any timestamped multimodal data could fit this paradigm.
Limitations and Unknowns
---
Sources