An Engineer's Annoyance
Suppose you are a backend engineer asked to integrate a payment gateway. The SDK docs list a dozen APIs, each with a dozen parameters. Your workflow looks like this:
1. Read the API docs 2. Fill in a JSON request body 3. Send the request 4. Read the response 5. If it fails, fill in another JSON based on the error code 6. Send again 7. Repeat until success
Every step is a "round trip." Chain a dozen APIs together and you get dozens of round trips. Latency accumulates, errors accumulate, and so does your impatience.
Now switch paradigms: write a Python script that chains all the API calls together. One execution, all calls completed in the same process, with parallelism, error handling, and branching all controlled in code. Latency drops from "dozens of round trips" to "one execution."
Which is faster? The answer is obvious. Yet it wasn't until August 2026, when Ishan Patel et al. published "The Bitter Lesson of Tool Calling" on arXiv, that someone systematically quantified the obvious.
Two Tool-Calling Paradigms
The paper compares two ways of calling tools:
JSON Tool Calling (Native)
The current mainstream. In each turn, the LLM outputs a JSON structure describing which tool to call and with what arguments. The system executes the call, feeds the result back into the context, and the LLM decides the next step.
Every step is a full LLM inference + one external call + one context update. N tool calls = N round trips.
Programmatic Tool Calling (PTC)
The approach the paper advocates. Instead of emitting JSON, the LLM directly writes a Python script. The script can import tool functions, loop, call in parallel, and handle exceptions. N tool calls = 1 script execution.
The key difference: in JSON mode, the LLM is a dispatcher—every call returns to the LLM; in PTC mode, the LLM is a programmer—it writes the code and steps away, leaving execution to the interpreter.
This looks like a mere engineering optimization, but it reveals a deeper structural problem.
A 14-Model Bake-Off on BFCL v4
The paper evaluates 14 models on BFCL (Berkeley Function Calling Leaderboard) v4, spanning GPT-4o through the GPT-5.6 Luna/Sol/Terra family, and Claude Sonnet 4.5 through Claude Sonnet 5.
Main Result
PTC matches or exceeds JSON tool calling on 11 of 14 models. The GPT-5.6 family improves by 10.6% on average.
More notable is the evolutionary trend:
- GPT-4o era: PTC trails JSON by 26.9% — models couldn't yet write code
- GPT-5-nano: PTC catches up to JSON for the first time — an inflection point
- GPT-5.6 era: PTC fully converges and starts to overtake
- JSON tool calling is the engineered approach — humans designed the JSON schemas, parameter types, and calling protocols within which the model operates
- PTC is the general compute approach — let the model write code; code is computation, and computation is general
- N calls = N inferences = N latencies
- All call state lives in the LLM context → context bloat
- Context bloat → context rot → lost calls
- N calls = 1 script execution = 1 latency
- Call state lives in script variables → no context bloat
- No bloat → no rot → no lost calls
- Euclid-MCP: let the LLM be the poet, let Prolog be the accountant. The LLM doesn't reason; reasoning is outsourced to a specialized tool. Same spirit as PTC: the LLM doesn't dispatch; dispatch is outsourced to the interpreter.
- colibrì: 1300 lines of C running 744-billion-parameter models on a 25GB laptop with no GPU. MoE decouples parameter count from memory requirements by 15x. Same spirit: decouple the number of tool calls from the number of LLM inferences.
- Rebucca: small model pre-screens, large model reviews. Division of labor beats unification. Same spirit: LLM writes code + interpreter executes code.
The authors name this phenomenon "The Bitter Lesson of Tool Calling" — an homage to Rich Sutton's classic. Sutton's core thesis: compute-based general methods eventually beat methods built on human-engineered knowledge. Here: letting models write code (general computation) eventually beats making models fill in JSON (engineered structure).
Three Structural Ablations
The paper designs three ablations, each exposing a structural issue.
1. Chaining
Chaining is: call A → get result → call B → get result → call C. In JSON mode, every step returns to the LLM; in PTC mode, one script handles it all.
PTC cuts wall-clock time for chained calls roughly in half on 13/14 models, with per-item latency ratios from 0.5 to 0.9. This isn't an optimization — it's a paradigm difference: N round trips vs 1 execution.
2. Parallel Fan-Out
Fan-out is calling N independent tools simultaneously. In JSON mode, the LLM must maintain state for N calls in its context. The paper finds a hard structural limit:
JSON tool calling completely loses tool calls at N=70-72 (Claude Sonnet 5) — not errors, just no calls made. The context grows too long and the model starts "forgetting" which tools to call.
PTC maintains 100% enumeration accuracy at N=100, because in a script it's just a for loop — larger N is just more iterations.
PTC matches or exceeds JSON fan-out on 13/14 models.
A takeaway worth remembering: JSON mode has a hard structural ceiling; PTC does not. The ceiling's position varies with model capability (GPT-4o's is lower than Claude Sonnet 5's), but the ceiling itself is structural — as long as "every call returns to the LLM," a context-length limit is inevitable.
3. Context Rot
Context rot: in long conversations, results from earlier tool calls "rot" — the model starts misreading, ignoring, or confusing old context.
JSON tool calling degrades by an average of 2.3% under context-rot conditions. PTC stays stable.
Why? In PTC mode, old call results don't need to stay in the LLM context — they live in the script's local variables. The LLM context only holds the script itself, not each call's result. This is a structural advantage: PTC is naturally resistant to context rot.
Why "Bitter Lesson"
Sutton's Bitter Lesson argues that general methods based on search and learning ultimately beat human-engineered approaches — a claim repeatedly validated in Go (AlphaGo), translation (NMT), and dialogue (GPT).
This paper applies the lesson to tool calling:
The evolutionary data is persuasive: in the GPT-4o era, PTC underperformed JSON because models couldn't write code; in the GPT-5.6 era, PTC overtakes because models can.
This matches Sutton's prediction exactly: once models are capable enough, general methods overtake engineered ones. JSON mode was a compromise for weak models — since models couldn't write code, humans designed JSON schemas to help them. Once models can write code, that compromise becomes a liability.
A Deeper Structural Insight
The paper's most memorable contribution isn't a specific number but a structural distinction:
In JSON mode, the LLM is a dispatcher — every tool call returns to the LLM. That means:
In PTC mode, the LLM is a programmer — it writes the code and steps away. That means:
This is a paradigm-level difference, not an optimization. The dispatcher paradigm has a structural ceiling; the programmer paradigm does not.
It evokes a cross-domain isomorphism: the octopus's "DNA pre-training + RNA inference-time computation". The octopus genome doesn't encode each synaptic connection; it encodes construction rules, with specific connections adjusted in real time via RNA editing in response to environmental stimuli. DNA is pre-training (writing rules), RNA is inference-time computation (executing rules).
PTC has the same structure: LLM pre-training (writes the script) + interpreter inference-time computation (executes the script). The LLM doesn't participate in every call, just as DNA doesn't participate in forming every synapse.
The isomorphism points to a general engineering principle: division of labor beats unification. Letting the LLM do what it's good at (writing code) and the interpreter do what it's good at (executing code) beats making the LLM do everything (writing JSON and dispatching execution).
An Honest Assessment
The paper has limitations.
First, PTC demands strong coding ability. PTC underperforming JSON in the GPT-4o era shows this paradigm has a threshold. For small models or weak coders, JSON mode may still be the better choice.
Second, PTC's security is an open question. Letting an LLM write and execute scripts means it can do anything code can do. JSON mode has natural structured constraints (schemas limit callable tools and parameters); PTC has none. Sandboxing, permissions, and auditing all need redesign.
Third, BFCL v4 is one specific benchmark. Whether PTC's advantage generalizes to other tool-calling scenarios (multi-turn dialogue, tool composition, error recovery) needs more validation.
But these don't diminish the core contribution: the paper systematically quantifies PTC vs JSON, exposes JSON mode's structural ceiling, and maps out how the Bitter Lesson plays out in tool calling.
Cross-Paper Resonance
This paper resonates with several other recent works:
Closing
The paper's most memorable line: when model capability is sufficient, general methods overtake engineered ones. That's Sutton's Bitter Lesson cashing out in tool calling.
JSON tool calling was a compromise for when models couldn't code. Once models can write code, that compromise becomes a structural ceiling. PTC isn't an optimization — it's a paradigm shift, from "LLM as dispatcher" to "LLM as programmer."
The consequences of this shift are only beginning to show. When LLMs no longer participate in every call, the cost, latency, and reliability of tool calling will all be redefined. And engineering efforts still piling onto JSON schemas may be repeating the bitter lesson Sutton warned us about.
---
Paper: https://arxiv.org/abs/2608.06370 HTML full text: https://arxiv.org/html/2608.06370v1 Related project: https://github.com/cameronking4/programmatic-tool-calling-ai-sdk