4-Hour Rust SQLite Rewrite, 80% Is Just Halftime: Cursor Turns AI Coding Into a Company
Source: Cursor official blog / sqllogictest / cursor/minisqlite
- Cursor official: https://cursor.com/blog/agent-swarm-model-economics
- SQLite sqllogictest: https://www.sqlite.org/sqllogictest/doc/trunk/about.wiki
- Public repo: https://github.com/cursor/minisqlite
This is not "just running more Claude Code instances." Cursor rebuilt the organizational structure: the strongest model acts as the planner, only splitting tasks, setting boundaries, and resolving conflicts; faster, cheaper models act as executors, each focused on one narrow problem at a time. Planners don't write implementations; executors don't change the overall design. Context is partitioned by responsibility—like a company, not a group of people fighting over the same code in one group chat.
The Real Differentiator Is Coordination Cost
Cursor ran four configurations in parallel: GPT-5.5 full-stack, Grok 4.5 full-stack, Opus 4.8 planning + Composer 2.5 execution, and Fable 5 planning + Composer 2.5 execution. At the four-hour mark, test pass rates for the new architecture ranged from 73% to 85%; all four eventually reached 100%.
The old architecture failed in concrete ways. The old Grok 4.5 version produced 68,000 commits and over 70,000 merge conflicts within two hours, and split into 54 crates including three mutually incompatible SQL packages. The new architecture fixed itself at 9 crates early on, with fewer than 1,000 cumulative conflicts over four hours.
Rather than pushing Git harder, Cursor built a purpose-built version control system from scratch. Early browser experiments peaked at about 1,000 commits per hour on Git; the new system's design peak is about 1,000 commits per second. Throughput isn't the point—the point is that conflict adjudication, design document references, and merge queues are all baked into the collaboration infrastructure.
Failure modes were documented candidly: two planners making conflicting designs, executors fighting over hot files, and code ossifying into a core nobody dares touch. Cursor's patches include shared design documents, compile-time references, third-party merge agents, and "permission to break old interfaces with justification." This part is worth more than the 80% result. It shows the main challenge in multi-agent engineering isn't whether models can write code—it's whether the organization stays under control.
A Model's Value Depends on the Seat It Occupies
Four-hour run costs ranged from $1,339 to $10,565. The most expensive GPT-5.5 full-stack setup spent $9,373 on executors alone; when Opus 4.8 planned and Composer 2.5 executed, the entire executor fleet cost just $411.
Executors consumed at least 69% of tokens—over 90% in most configurations. But planners, despite low token counts, carry high unit prices: in the Opus hybrid group, planners still accounted for roughly two-thirds of the cost. The bill yields a practical conclusion: frontier models don't need to blanket the entire pipeline—they just need to hold the high-ambiguity, high-impact decision points.
The public cursor/minisqlite repo is no longer just a four-hour demo. The README shows it has since grown to roughly 200,000 lines of Rust across 14 crates and 5,650 tests, with bidirectional read/write of SQLite format 3 files, plus WAL, transactions, triggers, window functions, and foreign keys. The boundaries are equally clear: no C API, no prepared statements, no cross-process concurrency, and no cost estimation based on sqlite_stat1. It is not yet an SQLite replacement.
Three Things I'll Be Watching
First, the 80% figure came from a company self-test—the questions, runtime environment, and human review were all controlled by Cursor. The code is public; independent reproduction is the next step.
Second, once agent swarms make "writing code" cheap, specifications become the new scarcity. The 835-page manual was rigorous enough for the swarm to compile text layer by layer into a system; bad requirements only get amplified in parallel.
Third, Cursor has pushed the competition from the model layer to the organization layer. Comparing Claude Code, Codex, and Cursor purely on SWE-bench is becoming less interesting. Whoever can keep dozens of agents from fighting, make them auditable and rollback-able is the one truly approaching an automated software company.
One-line takeaway: the most important number in this experiment isn't 80%—it's 70,000 conflicts compressed to under 1,000. The next step in AI coding isn't finding a smarter programmer model; it's building a corporate governance system that doesn't collapse for a crowd of models.