English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

code-review-graph: A Persistent Code Map That Cuts AI Coding Tool Context Costs by 100x

Forum topic · ✨步子哥 · 2026-08-06

Summary

AI coding assistants like Cursor and Claude Code re-scan the entire repository on every conversation, burning tens of thousands of tokens just to understand context. code-review-graph, an open-source tool by tirth8205, addresses this by building a persistent structural map of the codebase. It uses Tree-sitter for language-agnostic parsing (Python, TypeScript, Rust, Go), stores code relationships—function calls, class inheritance, module dependencies—as a graph in local SQLite, and exposes it via an MCP (Model Context Protocol) server that any compatible AI tool can query. This shifts context retrieval from search-driven full-repo scans (200-500ms, 50k+ tokens) to millisecond-level graph queries (1-5ms, a few hundred tokens)—a roughly 100x improvement from better granularity rather than faster engines. Unlike RAG-based semantic retrieval, it answers precise structural facts: who calls this function, what does this module depend on. The local-first SQLite design also beats cloud solutions like Sourcegraph on latency. Limitations include Tree-sitter language coverage, inability to capture semantic-level patterns, and early-stage incremental updates for large monorepos. GitHub: https://github.com/tirth8205/code-review-graph

code-review-graph: A Persistent Code Map for AI Coding Tools

When you ask Cursor to modify a function, it spends 30 seconds scanning the whole repository. Ask it to change a neighboring function, and it scans again for another 30 seconds. Same repo, same structural information—read from scratch every time.

This isn't because AI isn't smart enough. It's because AI has no map.

code-review-graph solves exactly this: it builds a persistent structural map of your codebase so AI tools only read the small piece they actually need.

The Problem: AI Coding Tools Have "Goldfish Memory"

Current AI coding tools (Cursor, Claude Code, Copilot) share a hidden cost: they start understanding the codebase from zero on every conversation.

The typical flow: you ask "where is this function called," and the tool either greps the whole repo (slow and noisy) or uses vector retrieval to find similar text fragments (potentially missing exact call relationships). Either way, every conversation re-scans, and every scan stuffs results into the context window.

A 100k-line codebase can cost an AI tool 50,000+ tokens per conversation just to "understand context." It's like having to relearn your way around the office every time you walk in—where the door is, where your desk is, where the meeting room is.

The root problem is a granularity mismatch: AI tools work at "file + line" granularity, but code structure lives at "function + call relationship + module dependency" granularity. Understanding code with the wrong granularity is like looking at a map through a microscope—enough magnification, but you'll never see the whole road.

The Approach: Build the Map First, Then Ask for Directions

code-review-graph works in three steps:

Step 1: Structured parsing with Tree-sitter. Tree-sitter is GitHub's open-source incremental parsing library that can parse source code into syntax trees in milliseconds. Instead of depending on language-specific LSPs (Language Server Protocol), code-review-graph uses Tree-sitter to uniformly handle all languages—Python, TypeScript, Rust, and Go all go through the same pipeline.

Step 2: Store the parse results as a graph. Not a file tree, not a symbol table, but a relationship graph: function A calls function B, class C inherits from class D, module E depends on module F. This graph is persisted in local SQLite. When an AI tool asks "who calls this function," it queries the graph instead of scanning the repo.

Step 3: Expose it to AI tools via MCP. MCP (Model Context Protocol) is Anthropic's standard protocol for letting AI tools access external data sources uniformly. code-review-graph implements an MCP server, so any MCP-compatible AI tool (Claude Code, Cursor, etc.) can query the code graph directly.

Why a "Map" Beats "Search"

Traditional AI tools are search-driven: ask a question, search for relevant files, stuff results into context. Every search starts from zero.

code-review-graph is map-driven: the structural graph is pre-built, so questions become graph queries instead of text searches.

| Dimension | Search-driven | Map-driven | |------|---------|---------| | First query | Full-repo scan | Graph query (ms-level) | | Subsequent queries | Re-scan | Graph query (ms-level) | | Context consumption | 50k+ tokens | A few hundred tokens | | Call relationships | Fuzzy grep matching | Exact graph edges | | Incremental updates | None | Only changed parts |

Key numbers: SQLite graph queries typically return in 1-5ms, while a traditional full-repo grep on a 100k-line codebase takes 200-500ms. The 100x speed difference doesn't come from a faster engine—it comes from smarter granularity.

This Is Not RAG

Isn't this just code-flavored RAG (Retrieval-Augmented Generation)?

No. RAG uses vector retrieval to find "semantically similar" text fragments, but code call relationships aren't captured by semantic similarity. Function A calling function B may involve two functions with completely different names and semantics—yet the call relationship is an exact structural fact.

RAG's core assumption—"semantic similarity ≈ relevance"—doesn't hold in rule-based worlds. As the Euclid-MCP experiment demonstrated, a 480B-parameter model can be just as bad as an 8B model at rule-based reasoning over 1,000 facts. A code call graph is exactly such a rule world: A calls B is a fact, not a probability.

code-review-graph does no semantic retrieval—it does structural queries. It answers not "which files are relevant to this question" but "who calls this function, what does it call, what modules does it depend on." Exact facts, no probabilities needed.

"Local-first" Isn't Decoration

code-review-graph emphasizes "local-first"—all data lives in local SQLite, nothing uploaded to the cloud. This isn't just about privacy; it's about performance.

Cloud solutions (like Sourcegraph) have 100-500ms latency; local SQLite queries run in 1-5ms. For AI coding tools, 100ms of query latency means 0.1 extra seconds per conversation—minutes of accumulated waiting per day. Local-first isn't a "privacy-friendly" marketing line; it's an engineering choice that's "100x faster."

Conceptual Lineage: From "Brute-Force Search" to "Structural Maps"

code-review-graph's core insight fits into a broader conceptual lineage:

  • Octopus RNA editing: DNA as pretraining + RNA as inference-time computation—change the blueprint, not the construction drawing
  • Slime mold externalized memory: memory externalized in slime trails, no neurons needed
  • Avian quantum magnetoreception: amplifier and sensor kept separate—division of labor beats unification
  • Euclid-MCP reasoning outsourcing: let the LLM be the poet, let Prolog be the accountant
  • code-review-graph structural map: externalize code structure into a persistent map—AI tools query instead of scan
  • All five point to the same principle: don't do the same thing harder—solve the problem at a different level. AI tools don't need stronger search; they need a pre-built map.

    Who Should Use It

  • Cursor/Claude Code users: If your codebase exceeds 10k lines, per-conversation context scanning costs are already significant. code-review-graph can cut context consumption by an order of magnitude.
  • MCP ecosystem participants: A benchmark implementation of the MCP protocol in the code-intelligence domain—worth studying for its interface design.
  • Codebase maintainers: Even without AI tools, the structural graph helps you understand dependencies in large codebases.
  • Limitations

  • Only supports languages Tree-sitter can parse: niche languages may be unsupported
  • Graph quality depends on Tree-sitter syntax trees: complex semantic-level relationships (like design patterns) can't be captured
  • Incremental updates are still early-stage: performance on large monorepos is unverified

Closing

The next breakthrough in AI coding tools won't come from bigger models, but from better context engineering. code-review-graph upgrades "understanding code" from "re-searching every time" to "querying a persistent map"—the same way human programmers work. You don't relearn your office layout every time you walk in; you have an internalized map.

A map is 100x faster than search—not because the map has a stronger engine, but because a map doesn't need to be redrawn every time.

---

GitHub: https://github.com/tirth8205/code-review-graph Website: https://code-review-graph.com MCP Protocol: https://modelcontextprotocol.io

Tags

#code-review-graph#mcp#tree-sitter#ai-coding-tools#context-engineering#code-analysis#sqlite#local-first

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603051