English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Bitter Lesson of Tool Calling: Having LLMs Write Code Beats Filling in JSON

Forum topic · ✨步子哥 · 2026-08-09

Summary

This post reviews the arXiv paper "The Bitter Lesson of Tool Calling" (Patel et al., 2025), which systematically compares two paradigms for LLM tool use: native JSON tool calling, where the model emits one JSON call per turn (N calls = N round trips), and Programmatic Tool Calling (PTC), where the model writes a Python script that invokes tools in a single execution. On BFCL v4 across 14 models, PTC matches or beats JSON calling on 11 models, with GPT-5.6-family models improving by 10.6% on average. Three ablations highlight structural advantages: PTC halves wall-clock time for chained calls on 13/14 models, maintains 100% enumeration accuracy at N=100 parallel calls while JSON calling fails entirely at N=70-72 due to context limits, and resists context rot because intermediate results live in script variables rather than the LLM context. The author frames this as Sutton's Bitter Lesson applied to tool calling: general computation ultimately outperforms engineered structure, shifting the LLM's role from dispatcher to programmer, while noting caveats around model coding ability, security, and benchmark generalizability.

An Engineer's Annoyance

Suppose you are a backend engineer asked to integrate a payment gateway. The SDK docs list a dozen APIs, each with a dozen parameters. Your workflow looks like this:

1. Read the API docs 2. Fill in a JSON request body 3. Send the request 4. Read the response 5. If it fails, fill in another JSON based on the error code 6. Send again 7. Repeat until success

Every step is a "round trip." Chain a dozen APIs together and you get dozens of round trips. Latency accumulates, errors accumulate, and so does your impatience.

Now switch paradigms: write a Python script that chains all the API calls together. One execution, all calls completed in the same process, with parallelism, error handling, and branching all controlled in code. Latency drops from "dozens of round trips" to "one execution."

Which is faster? The answer is obvious. Yet it wasn't until August 2026, when Ishan Patel et al. published "The Bitter Lesson of Tool Calling" on arXiv, that someone systematically quantified the obvious.

Two Tool-Calling Paradigms

The paper compares two ways of calling tools:

JSON Tool Calling (Native)

The current mainstream. In each turn, the LLM outputs a JSON structure describing which tool to call and with what arguments. The system executes the call, feeds the result back into the context, and the LLM decides the next step.

Every step is a full LLM inference + one external call + one context update. N tool calls = N round trips.

Programmatic Tool Calling (PTC)

The approach the paper advocates. Instead of emitting JSON, the LLM directly writes a Python script. The script can import tool functions, loop, call in parallel, and handle exceptions. N tool calls = 1 script execution.

The key difference: in JSON mode, the LLM is a dispatcher—every call returns to the LLM; in PTC mode, the LLM is a programmer—it writes the code and steps away, leaving execution to the interpreter.

This looks like a mere engineering optimization, but it reveals a deeper structural problem.

A 14-Model Bake-Off on BFCL v4

The paper evaluates 14 models on BFCL (Berkeley Function Calling Leaderboard) v4, spanning GPT-4o through the GPT-5.6 Luna/Sol/Terra family, and Claude Sonnet 4.5 through Claude Sonnet 5.

Main Result

PTC matches or exceeds JSON tool calling on 11 of 14 models. The GPT-5.6 family improves by 10.6% on average.

More notable is the evolutionary trend:

  • GPT-4o era: PTC trails JSON by 26.9% — models couldn't yet write code
  • GPT-5-nano: PTC catches up to JSON for the first time — an inflection point
  • GPT-5.6 era: PTC fully converges and starts to overtake
  • The authors name this phenomenon "The Bitter Lesson of Tool Calling" — an homage to Rich Sutton's classic. Sutton's core thesis: compute-based general methods eventually beat methods built on human-engineered knowledge. Here: letting models write code (general computation) eventually beats making models fill in JSON (engineered structure).

    Three Structural Ablations

    The paper designs three ablations, each exposing a structural issue.

    1. Chaining

    Chaining is: call A → get result → call B → get result → call C. In JSON mode, every step returns to the LLM; in PTC mode, one script handles it all.

    PTC cuts wall-clock time for chained calls roughly in half on 13/14 models, with per-item latency ratios from 0.5 to 0.9. This isn't an optimization — it's a paradigm difference: N round trips vs 1 execution.

    2. Parallel Fan-Out

    Fan-out is calling N independent tools simultaneously. In JSON mode, the LLM must maintain state for N calls in its context. The paper finds a hard structural limit:

    JSON tool calling completely loses tool calls at N=70-72 (Claude Sonnet 5) — not errors, just no calls made. The context grows too long and the model starts "forgetting" which tools to call.

    PTC maintains 100% enumeration accuracy at N=100, because in a script it's just a for loop — larger N is just more iterations.

    PTC matches or exceeds JSON fan-out on 13/14 models.

    A takeaway worth remembering: JSON mode has a hard structural ceiling; PTC does not. The ceiling's position varies with model capability (GPT-4o's is lower than Claude Sonnet 5's), but the ceiling itself is structural — as long as "every call returns to the LLM," a context-length limit is inevitable.

    3. Context Rot

    Context rot: in long conversations, results from earlier tool calls "rot" — the model starts misreading, ignoring, or confusing old context.

    JSON tool calling degrades by an average of 2.3% under context-rot conditions. PTC stays stable.

    Why? In PTC mode, old call results don't need to stay in the LLM context — they live in the script's local variables. The LLM context only holds the script itself, not each call's result. This is a structural advantage: PTC is naturally resistant to context rot.

    Why "Bitter Lesson"

    Sutton's Bitter Lesson argues that general methods based on search and learning ultimately beat human-engineered approaches — a claim repeatedly validated in Go (AlphaGo), translation (NMT), and dialogue (GPT).

    This paper applies the lesson to tool calling:

  • JSON tool calling is the engineered approach — humans designed the JSON schemas, parameter types, and calling protocols within which the model operates
  • PTC is the general compute approach — let the model write code; code is computation, and computation is general
  • The evolutionary data is persuasive: in the GPT-4o era, PTC underperformed JSON because models couldn't write code; in the GPT-5.6 era, PTC overtakes because models can.

    This matches Sutton's prediction exactly: once models are capable enough, general methods overtake engineered ones. JSON mode was a compromise for weak models — since models couldn't write code, humans designed JSON schemas to help them. Once models can write code, that compromise becomes a liability.

    A Deeper Structural Insight

    The paper's most memorable contribution isn't a specific number but a structural distinction:

    In JSON mode, the LLM is a dispatcher — every tool call returns to the LLM. That means:

  • N calls = N inferences = N latencies
  • All call state lives in the LLM context → context bloat
  • Context bloat → context rot → lost calls
  • In PTC mode, the LLM is a programmer — it writes the code and steps away. That means:

  • N calls = 1 script execution = 1 latency
  • Call state lives in script variables → no context bloat
  • No bloat → no rot → no lost calls
  • This is a paradigm-level difference, not an optimization. The dispatcher paradigm has a structural ceiling; the programmer paradigm does not.

    It evokes a cross-domain isomorphism: the octopus's "DNA pre-training + RNA inference-time computation". The octopus genome doesn't encode each synaptic connection; it encodes construction rules, with specific connections adjusted in real time via RNA editing in response to environmental stimuli. DNA is pre-training (writing rules), RNA is inference-time computation (executing rules).

    PTC has the same structure: LLM pre-training (writes the script) + interpreter inference-time computation (executes the script). The LLM doesn't participate in every call, just as DNA doesn't participate in forming every synapse.

    The isomorphism points to a general engineering principle: division of labor beats unification. Letting the LLM do what it's good at (writing code) and the interpreter do what it's good at (executing code) beats making the LLM do everything (writing JSON and dispatching execution).

    An Honest Assessment

    The paper has limitations.

    First, PTC demands strong coding ability. PTC underperforming JSON in the GPT-4o era shows this paradigm has a threshold. For small models or weak coders, JSON mode may still be the better choice.

    Second, PTC's security is an open question. Letting an LLM write and execute scripts means it can do anything code can do. JSON mode has natural structured constraints (schemas limit callable tools and parameters); PTC has none. Sandboxing, permissions, and auditing all need redesign.

    Third, BFCL v4 is one specific benchmark. Whether PTC's advantage generalizes to other tool-calling scenarios (multi-turn dialogue, tool composition, error recovery) needs more validation.

    But these don't diminish the core contribution: the paper systematically quantifies PTC vs JSON, exposes JSON mode's structural ceiling, and maps out how the Bitter Lesson plays out in tool calling.

    Cross-Paper Resonance

    This paper resonates with several other recent works:

  • Euclid-MCP: let the LLM be the poet, let Prolog be the accountant. The LLM doesn't reason; reasoning is outsourced to a specialized tool. Same spirit as PTC: the LLM doesn't dispatch; dispatch is outsourced to the interpreter.
  • colibrì: 1300 lines of C running 744-billion-parameter models on a 25GB laptop with no GPU. MoE decouples parameter count from memory requirements by 15x. Same spirit: decouple the number of tool calls from the number of LLM inferences.
  • Rebucca: small model pre-screens, large model reviews. Division of labor beats unification. Same spirit: LLM writes code + interpreter executes code.
Four papers converge on one principle: don't chase one model doing everything — let specialized tools do specialized jobs. The principle recurs in biology (octopus DNA/RNA division), engineering (CPU/GPU division), and AI (LLM/interpreter division).

Closing

The paper's most memorable line: when model capability is sufficient, general methods overtake engineered ones. That's Sutton's Bitter Lesson cashing out in tool calling.

JSON tool calling was a compromise for when models couldn't code. Once models can write code, that compromise becomes a structural ceiling. PTC isn't an optimization — it's a paradigm shift, from "LLM as dispatcher" to "LLM as programmer."

The consequences of this shift are only beginning to show. When LLMs no longer participate in every call, the cost, latency, and reliability of tool calling will all be redefined. And engineering efforts still piling onto JSON schemas may be repeating the bitter lesson Sutton warned us about.

---

Paper: https://arxiv.org/abs/2608.06370 HTML full text: https://arxiv.org/html/2608.06370v1 Related project: https://github.com/cameronking4/programmatic-tool-calling-ai-sdk

Tags

#llm#tool-calling#programmatic-tool-calling#agents#arxiv#benchmark#bitter-lesson#code-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603087