Self-Speculating Agents: Eliminating Tool-Call Dead Waiting — Nobody Knows You Better Than Yourself
This post reviews the paper *Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent–Speculator RL* (Ji et al., 2026, arXiv:2607.25816, UC Santa Barbara + LinkedIn).
Key points
- The agent bottleneck is not thinking — it's waiting. Token generation is fast; wall-clock time is dominated by network I/O on each tool call (0.5–2s per round), citing Nichols et al. (2025) and Hooper et al. (2026).
- The Speculator-Agent Gap. External small draft models used for tool-call speculation poorly predict the agent's actual decisions. On MuSiQue and τ-bench (Hit@1 = exact match of tool name + arguments):
- Self-speculating agent: one model, two modes. An *Agent mode* (reason → tool call → observe → ...) and a *Speculator mode* that takes an intermediate trajectory plus a fixed speculation suffix and predicts the next tool call. Both modes share the prefix KV cache, so speculation costs no extra weights and no extra cache. The fixed suffix contains no task-specific information, only a mode switch.
- Joint Agent-Speculator RL. Naively training prediction hurts task success. The method: sample on-policy rollouts, construct speculation queries from them, then alternate optimization (agent updates vs. speculator updates) with three stabilizers: SFT warmup, optimizer resets on mode switch, and a 4-step-agent / 8-step-speculator schedule.
- Read-only tools only; extending to writes may need dry-run or transactional rollback.
- Only tested at 4B scale (Qwen3-4B, Qwen3.5-4B); optimization dynamics may differ at larger scales.
- Only search QA and conversational tool calling tested; code execution, web navigation, and multi-agent settings remain unexplored.
- Open questions: predicting tool *choice* (not just arguments) with many tools, Hit@k via parallel speculation, and feeding predictions back to the agent to improve decision quality.
| Speculator | MuSiQue Hit@1 | τ-bench Hit@1 | Extra GPU memory | |:---|:---|:---|:---| | Qwen3-0.6B (external) | 14.7 | 25.4 | +2.01 GB | | Qwen3-1.7B (external) | 12.3 | 22.1 | +4.06 GB | | Qwen3-4B (itself) | 25.3 | 37.2 | +0 |
The agent itself is the best speculator of its own next call — two models, even from the same family, have different decision boundaries.
Main results
| Model | Stage | Avg Hit@1 | Avg task success | |:---|:---|:---|:---| | Qwen3-4B | Base | 29.1 | — | | Qwen3-4B | + SFT warmup | 44.1 | — | | Qwen3-4B | + Joint RL | 61.2 | 26.6 → 27.7 ↑ | | Qwen3.5-4B | Base | 33.5 | — | | Qwen3.5-4B | + SFT warmup | 48.9 | — | | Qwen3.5-4B | + Joint RL | 66.3 | 49.2 → 50.6 ↑ |
Hit@1 gains of ~39% (relative) with task success *improving*, not degrading.
Ablation on schedule
| Schedule | Avg Hit@1 | Avg success | |:---|:---|:---| | 1:1 | 31.8 | 10.3 (collapse) | | 2:2 | 41.8 | 22.9 | | 4:4 | 50.3 | 24.4 | | 4:8 | 55.2 | 26.1 |
Per-step switching destabilizes training; the speculator needs longer runs of consecutive updates to track the agent's current call distribution.
Cross-domain generalization
Trained on SearchQA, tested on τ-bench: Hit@1 improves 35.6 → 45.6, but task success drops 21.6 → 17.6. Prediction skill transfers across domains; task-solving skill does not.
Engineering insights
1. "Nobody knows you better than yourself" is math, not a platitude. A small model is not a shrunk copy of a large one — their learned decision paths differ. Self-speculation eliminates the gap because speculator and agent share one decision boundary: the model simply says earlier what it was going to say anyway. 2. Multi-objective training needs contiguous signals. The 1:1 collapse and 4:8 optimum generalize: alternate objectives need enough consecutive steps, and optimizer states (momentum) of different objectives should not interfere. 3. Read-only tools are the safety boundary. Speculation only applies to read-only tools (search, retrieval, queries). Pre-executing state-changing tools (orders, updates, messages) risks side effects — you may pre-search, but you may not pre-order.
Architectural reflection: unification vs. division of labor
The author contrasts this with prior "division beats unification" designs (LLM + Prolog, small-model screening + large-model review). Here unification wins because the speculator's *task is to predict the agent*, which requires a shared decision boundary. The refined principle: unify when two tasks must share a decision boundary; divide when they need different capabilities.
Limitations and open questions
---
Paper: Ji et al. (2026). *Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent–Speculator RL.* arXiv:2607.25816. https://arxiv.org/abs/2607.25816
Authors: Jiabao Ji, Yujian Liu, Li An, Rohit Jain, Gungor Polatkan, Siyu Zhu, Shiyu Chang (UC Santa Barbara + LinkedIn Inc.)