RynnValue: Time-Distance Supervision for Robot Value Models Over 7,000 Hours of Data
> Indexed by AI HOT (arXiv 2608.09853, HuggingFace Daily Papers, 2026-08-10). Secondary sources: arXiv cs.RO paper page and HuggingFace paper page.
1. The Problem: Where Robot Reward Models Are Stuck
In generalist robot policy learning, value models are caught between two contradictions:
- Supervision signals are hard to transfer: Existing methods anchor supervision within tasks — preferences or normalized progress. These anchors largely fail when moving across embodiments or data sources.
- Annotation costs don't scale: Preference labeling requires humans to compare two trajectories; progress labeling requires humans to define stage boundaries. How long would it take humans to annotate 7,000 hours of data? Nobody can compute it.
- Labels are naturally generatable: Every episode carries timestamps; cost-to-go is computed directly from them with no human effort.
- Cross-embodiment by construction: Time distance does not depend on embodiment-specific features.
- Zero-shot transfer: Because the supervision signal is decoupled from embodiment, the learned value function generalizes directly to unseen tasks, unknown embodiments, and novel viewpoints.
- Training data: 7,000+ hours, ~3 million instruction-conditioned episodes, no preference or progress annotations
- Kendall's tau_a = 0.675: exceeds the fully preference-supervised SOTA (0.655)
- Improvement factor: over 2.3× the Kendall's tau of the progress-only baseline (0.292)
- Real-world policy success rates:
- Online learning: 52.5% → 72.5% (+20 percentage points)
- Offline learning: 63.8% → 82.5% (+18.7 percentage points)
- Zero-shot transfer: unseen tasks / unseen embodiments / novel viewpoints
- Plugs into any RL algorithm: PPO / SAC / IQL all work, since the interface is a standard dense reward.
- Composes with VLA models: RynnValue as critic, π0 / OpenVLA as actor, with a clean division of labor.
- Drives sim-to-real: run the time-distance value model in simulation; no retraining for real-robot deployment.
- Open-weight release: whether weights ship with the paper determines industrial adoption speed.
- Cross-vendor training: if teams like NVIDIA GR00T, Google DeepMind RT-2, and Physical Intelligence π0 adopt the same time-distance route as a control, value modeling will converge fast.
- Coupling with VLA models measured: the +20 percentage point gain for the value model alone is clear; how much it adds inside OpenVLA / π0 / Helix is the real question.
RynnValue's (arXiv 2608.09853, 23 pages, 5 figures) solution is blunt: drop preferences, drop progress, and use timestamps directly as labels.
2. Why Time Distance Works as a Supervision Signal
The core idea — treat the directed cost-to-go from each observation to a language-specified goal within each trajectory as the label — is almost too simple to work:
But there is a hidden failure mode: the model may take a shortcut — reading elapsed time rather than task completion. RynnValue suppresses this with three techniques:
1. Random temporal sampling: shuffling trajectory order during training, forcing the model to read whole task structure instead of temporal anchors. 2. Temporal shuffling: further permuting segments within a task to strengthen progress discrimination. 3. Value-isolation attention: isolating the value-signal channel in attention layers so predictions are sensitive to failure and rollback rather than to the point in time.
3. Hard Numbers on Data and Benchmarks
RynnValue is evaluated on RBM-EVAL-OOD (an out-of-distribution robotic manipulation benchmark):
Beating a human-preference-annotated SOTA with timestamps alone is the key inversion — it suggests the bottleneck in embodied value modeling is not architecture but annotation cost.
4. What Shifting from Preferences to Time Means
Preference annotation has dominated embodied foundation models for the past two years — Collective Human Feedback-style work, RLHF-style preference pairs, Bradley-Terry models — all relying on humans comparing two trajectories.
RynnValue charts a completely different data acquisition path:
| Dimension | Preference labeling | Time distance | |---|---|---| | Label source | Human compares two trajectories | Timestamps already in trajectories | | Annotation cost | O(n²) pairing | O(1) direct generation | | Embodiment transfer | Poor (preferences are embodiment-specific) | Strong (time distance is embodiment-agnostic) | | Data scale ceiling | ~10k hours (human-limited) | Tens of millions of hours (auto-generated) | | Failure sensitivity | Strong (humans spot failures) | Medium (requires model design to suppress) |
The last row is RynnValue's engineering core — it compensates for the failure-sensitivity gap with attention design. The architecture is not complex; value-isolation attention is the key trick, forcibly separating the value-prediction channel from the "elapsed time shortcut."
5. Chain Effects of a Potential Reward Interface
RynnValue's engineering value is not just the base model but the dense reward it produces.
The paper converts the time-distance value function into a potential-based shaping reward — the RL-standard approach that provably preserves optimal policies (it changes paths, not solutions).
Chain effects of this interface:
This is in the same lineage as OpenAI o1-style reasoning models using process reward models as verifiers — generalizing reward signals is a key step toward industrializing embodied AI.
6. Limitations and Next Steps
RynnValue is not a silver bullet. Three real limitations:
1. Timestamp reliability: if source data timestamps are inaccurate (e.g., unsynchronized clocks across robot systems), labels carry noise. The paper assumes data from one device cluster, but industrial deployment spans vendors. 2. Task-structure stability: the value model assumes a clear "start-to-goal" task structure. If the goal itself is vague (e.g., "make the user satisfied"), the time-distance signal degenerates into duration prediction. 3. Embodiment OOD limits: zero-shot works on unseen embodiments, but cross-morphology transfer (legged → wheeled → aerial) is not thoroughly validated.
Key next steps:
Core numbers: 7,000+ hours of training data, 3 million instruction episodes, Kendall's tau_a 0.675 vs SOTA 0.655 vs progress-only 0.292, real-world success rate 52.5%→72.5% online and 63.8%→82.5% offline, 23 pages / 5 figures, submitted 2026-08-10. Timeline: arXiv 2608.09853 v1 (2026-08-11 01:09 UTC), cs.RO, indexed by HuggingFace Daily Papers. Sources: arXiv, HuggingFace Daily Papers, research discussion groups.