PTRM: A 7M-Parameter Tiny Recursive Model Goes Probabilistic — Noise Injection and Q-Head Reuse for Test-Time Scaling
Paper: Probabilistic Tiny Recursive Model (arXiv:2605.19943) Authors: Amin Sghaier, Ali Parviz, Alexia Jolicoeur-Martineau — Mila – Quebec AI Institute, ILLS & ETS Montreal
Key points
- Background: TRM (Tiny Recursive Model), introduced in October 2025 by Alexia Jolicoeur-Martineau (Samsung SAIL Montreal), is a 7M-parameter, two-layer network that recursively refines a latent reasoning state
zand an answery, trained with deep supervision. Reported results: Sudoku-Extreme 87.4%, Maze-Hard 85.3%, ARC-AGI-1 44.6%, ARC-AGI-2 7.8% — beating Gemini 2.5 Pro (37.0% / 4.9%) and DeepSeek R1 (15.8% / 1.3%) on ARC-AGI. - The flaw: TRM's recursion is fully deterministic. Given the same
(x, y, z), it always produces the same trajectory. If the initial guess lands in a bad latent basin, the model is stuck with no exploration mechanism. - PPBench puzzles: PTRM (7M) 91.2% at $0.001/puzzle vs Claude Opus 4.6 at 34.7% ($2.91/puzzle) and a hypothetical ensemble of 7 frontier LLMs with a perfect verifier at 55.1% ($38.51 per correct answer).
- Sudoku-Extreme: 98.75% vs TRM's 87.4% (+11.35 points).
- Width scaling (PPBench): K=1: 76.4% → K=10: 83.1% → K=50: 87.6% → K=100: 89.5% (+13.1 points).
- Verifier quality: best-Q@K stays within ~1 point of the pass@K upper bound on Sudoku and PPBench, confirming the Q-head is a highly reliable selector.
- Noise at each step (not just initialization) gives trajectories continuous chances to hop into better solution basins; moderate noise escapes local optima, excessive noise destroys existing structure.
- Width vs depth: depth scaling (more recursion steps) is serial and still deterministic; width scaling (K parallel trajectories) naturally parallelizes on GPUs and explores diverse basins.
- No retraining needed: the framework applies directly to already-trained TRM checkpoints, is task-agnostic, and turns a training-time byproduct into a test-time decision mechanism.
The PTRM method
PTRM makes two minimal changes:
1. Noise injection at every depth step: z_t = rec(x, z_{t-1} + ε, y_{t-1}), with ε ~ N(0, σ²I). Running K trajectories in parallel, each step perturbs the path, so trajectories accumulate different noise and can escape local optima.
2. Q-head reuse: TRM's Q-head (a correctness classifier used for adaptive compute time / early stopping during training) is normally discarded at inference. PTRM repurposes it as a test-time trajectory selector: decode K candidate answers, score each with the Q-head, and return the highest-scoring one (best-Q@K).
Hyperparameters: parallel trajectories K = 1–100, supervised depth D = 8–16, task-dependent noise scale σ.
Results
Task-dependent gains
| Task | pass@K gain | best-Q@K | Note | |---|---|---|---| | Sudoku-Extreme | +11.35 | ~98.75% | Q-head very effective | | PPBench | +13.1 | ~91.2% | Q-head very effective | | Maze-Hard | +11.83 | 86.73% | pass@K is 95.63% but Q-head can't pick it out (~10% verifier ceiling) | | ARC-AGI-2 | +1.11 | ~9% | Minimal benefit | | Heyawake | +1.11 | – | Minimal benefit |
Pattern: easily verifiable tasks see large gains; hard-to-verify tasks limit the approach because the Q-head loses discriminative power.
Why it works
Limitations and open questions
1. Evaluated only on reasoning puzzles (5/6 PPBench puzzle types, 9×9 / 10×10 grids); general tasks untested. 2. Verifier ceiling: on Maze-Hard and ARC-AGI-2 the Q-head is imperfect, capping gains. 3. Noise scale σ is chosen per task with no automated methodology — too small fails to escape, too large destroys solutions. 4. Cost curve: K=100 costs 100× inference compute (still trivial vs LLMs); marginal returns diminish but no ceiling seen up to K=100. K=1000 untested. 5. The Q-head's reliability holds on the training distribution; heavy noise or harder tasks may require external verifiers.
Takeaway
PTRM frames a third generation of tiny recursive reasoning (HRM → TRM → PTRM): a 7M-parameter model doesn't need to become a 700B model to reason well — it needs stochastic exploration in latent space, parallel trajectories, and its own learned self-verification to pick the best path, at a fraction of the cost.
Reference: arXiv:2605.19943 — https://arxiv.org/abs/2605.19943. Predecessor: *Less is More: Recursive Reasoning with Tiny Networks*.