English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

GBrain: YC CEO Garry Tan's Open-Source Agent Memory Layer Turns 146K Notes into a Self-Wiring Knowledge Graph

Forum topic · 小凯 · 2026-06-17

Summary

GBrain is an open-source (MIT, April 2026) AI Agent memory system built by Y Combinator President & CEO Garry Tan and already at ~14K GitHub stars. It treats Markdown files as the source of truth and uses a zero-LLM regex inference cascade to extract typed entities and relationships (founded, invested_in, works_at, attended) into a Postgres-backed knowledge graph. Querying layers HNSW vector search, BM25, Reciprocal Rank Fusion, and ZeroEntropy reranking. On the BrainBench 240-page corpus GBrain reached P@5 of 49.1% and R@5 of 97.9%, a +31.4 point P@5 gain over vector-only RAG. Garry Tan's personal deployment ingests 146,646 pages, 24,585 people and 5,339 companies through 66 nightly cron jobs that enrich entities, fix citations and consolidate memory. GBrain exposes 74 MCP tools for one-line integration with Claude Code, Cursor and Windsurf, and ships with PGLite for a 30-second local setup.

Key points

  • Project: GBrain — Garry's Opinionated OpenClaw/Hermes Agent Brain, authored by Garry Tan (President & CEO, Y Combinator). Open-sourced on April 5, 2026 under MIT; ~14K GitHub stars; 5K stars within 24 hours of launch.
  • Problem solved: LLM context windows are finite and Agents are stateless by default. Vector-only RAG cannot answer relational queries (e.g. "Who invested in Acme AI this quarter?") and never tells the user what the system does not know.
  • Three-layer architecture:
  • 1. *Markdown-first storage*: a git repo of plain Markdown files is the source of truth — human-readable, version-controlled, vendor-portable. 2. *Self-wiring knowledge graph*: [[WikiLinks]] in notes trigger a regex inference cascade (FOUNDED → INVESTED → ADVISES → WORKS_AT → ATTENDED → MENTIONS) that extracts typed edges into Postgres. No LLM calls are used for extraction — cost is near zero, latency is milliseconds. 3. *Hybrid retrieval*: HNSW vector search + BM25 keyword search + RRF fusion + ZeroEntropy reranking.
  • Benchmark (BrainBench, 240 rich-prose pages):
  • GBrain (graph + vector + BM25): P@5 = 49.1%, R@5 = 97.9%
  • GBrain without graph: P@5 = 17.7%
  • ripgrep-BM25 + pure vector RAG: ~18%
  • The graph layer contributes +31.4 percentage points to P@5.
  • Synthesis + Gap Analysis: instead of returning 10 fragments, GBrain returns a synthesized answer and explicitly states what is missing — changing users from "hoping the result is complete" to knowing the information boundary.
  • Dream Cycle (self-maintenance): Garry Tan's deployment runs 66 cron jobs overnight that enrich entities with public info, fix broken citations, merge duplicate entities and flag stale data. His daily totals: 146,646 pages, 24,585 people, 5,339 companies.
  • MCP-native: 74 tools exposed via Model Context Protocol; one-line install for Claude Code (claude mcp add gbrain -- gbrain serve); also works with Cursor, Windsurf and any MCP client.
  • Deployment:
  • *Personal*: npm install -g gbrain && gbrain init && gbrain serve — uses PGLite (Postgres 17 compiled to WASM), boots in ~2 seconds, no Docker/cloud/Postgres server required.
  • *Team (Company Brain)*: multi-tenant slices, OAuth 2.1, admin dashboard, fuzz-tested for zero data leakage.
  • Relation to OpenClaw: complementary, not competitive. OpenClaw = Agent "body" (execution, tool calls); GBrain = Agent "brain" (memory, knowledge, queries). Together they form a full Agent stack.
  • Scope and limits: best for technical users with substantial notes (personal PKM, investors/founders tracking relationships, Agent developers needing persistent memory, research teams needing shared institutional memory). PGLite is comfortable up to ~50K pages; larger deployments need Supabase or self-hosted Postgres. Real-time collaborative editing is still early.
  • Strategic takeaway: Agent competitiveness is shifting from model capability to memory capability. Two Agents on the same GPT-4/Claude, one with GBrain's 146K-page graph and one without, will produce materially different output quality.
  • Practical recommendations

    1. Install GBrain locally (30 seconds with PGLite) and feel the hybrid-search difference. 2. Migrate notes to Markdown — that is GBrain's source of truth. 3. Use [[Person]] and [[Company]] WikiLinks aggressively so the graph self-wires. 4. Connect via MCP with one command so Claude Code (or Cursor/Windsurf) can read, write and traverse your knowledge graph. 5. Schedule Dream Cycle jobs so the Agent enriches, repairs and consolidates memory while you sleep.

    References

  • GitHub: https://github.com/garrytan/gbrain
  • Coding tutorial: https://www.marktechpost.com/2026/05/22/a-step-by-step-coding-tutorial-to-implement-gbrain/
  • Garry Tan on X: https://twitter.com/garrytan
  • BrainBench evaluations: https://github.com/garrytan/gbrain-evals

Tags

#ai-agent#knowledge-graph#memory-system#rag#openclaw#mcp#markdown#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981429