Paper Overview
Field: Machine Learning Authors: Benhao Huang, Zhengyang Geng, Zico Kolter Published: 2025-05-20 arXiv: 2505.15988
Abstract (translated)
Scaling test-time compute by iteratively updating a latent state has emerged as a powerful paradigm for reasoning. Yet the internal mechanisms that enable these iterative models to generalize beyond memorized patterns remain unclear. We hypothesize that generalizable reasoning arises from learning task-conditioned attractors: latent dynamical systems whose stable fixed points correspond to valid solutions. We formalize this process through Equilibrium Reasoners (EqR), which enable test-time scaling without external verifiers or task-specific priors. EqR scales internal dynamics along two axes: depth, by running more iterations, and breadth, by aggregating stochastic trajectories from multiple initializations. Empirically, gains from test-time scaling are tightly coupled with stronger convergence toward solution-aligned attractors.
Key points
- Attractor hypothesis: Generalizable reasoning stems from learning task-conditioned attractors—latent dynamical systems whose stable fixed points correspond to valid solutions.
- Two scaling axes: EqR scales test-time compute via *depth* (more iterations) and *breadth* (aggregating stochastic trajectories from multiple initializations), requiring no external verifier or task-specific priors.
- Adaptive compute allocation: Gains from test-time scaling are tightly coupled with stronger convergence to solution-aligned attractors, letting the network adapt compute to task difficulty. Easy cases converge within 1–5 steps; hard cases benefit from large-scale scaling.
- Results: By unrolling up to an equivalent of 40,000 layers, scalable implicit reasoning improves feedforward model accuracy on Sudoku-Extreme from 2.6% to over 99%.
- Takeaway: Learned attractor landscapes provide a useful mechanistic lens for understanding scalable reasoning in iterative implicit models.
*Auto-collected on 2026-05-22*