GraphRAG Open-Source Projects Compared: Who's Swimming Naked, Who's Fighting Hard
If you have thousands of research reports, tens of thousands of support tickets, or an entire internal wiki, traditional RAG struggles: ask "what are the core themes of this dataset" or "how does Company A's supply chain change affect Company B's quarterly earnings," and it stalls.
Traditional RAG chunks documents, computes vector similarity, and always retrieves the "few most locally similar fragments." There are no connections between fragments, so cross-document reasoning is impossible and global summarization is out of reach.
GraphRAG's idea is simple: reweave scattered text fragments into a knowledge network. Entities are nodes, relationships are edges, communities are clusters. Retrieval no longer asks "which text is most similar" but walks along the graph. The catch: building that web is expensive — Microsoft's GraphRAG can burn a large chunk of your API quota on indexing alone. Hence a wave of open-source projects, some subtracting, some adding, some taking entirely different paths.
Key points
- Microsoft GraphRAG (31k+ stars): chunking → LLM entity/relation extraction → knowledge graph → Leiden hierarchical community detection → LLM community summaries → retrieval via community summaries. Three query modes: Local Search, Global Search (Map-Reduce over community summaries), and the newer DRIFT Search. Indexing is extremely expensive — one test on the first nine chapters of *Journey to the West* showed index build costs over 10x LightRAG's. Incremental indexing arrived in 0.5.0 but still requires community recomputation. LazyGraphRAG claims 0.1% of original indexing cost via "cluster first, summarize on demand," but isn't fully integrated yet. Best for high-value, low-frequency global analysis (financial research, literature reviews, trend scanning).
- LightRAG (29k+ stars, HKU HKUDS): builds a graph plus key-value indexes, embeds nodes and edges, and uses dual-level retrieval — Local extracts concrete keywords to retrieve nodes + one-hop neighbors; Global extracts abstract keywords to retrieve edges (relations) + endpoints; Hybrid and Mix (adding naive vector search) modes also exist. Its biggest selling point is incremental updates via set-union merging — no community recomputation. Benchmark repo NanGePlus/LightRAGTest reports 99.98% retrieval efficiency gain, ~12x lower latency, 99%+ fewer API calls per query, 12x+ lower token consumption. Multimodal support via RAG-Anything. Weakness: weaker complex reasoning and global integration (no community summaries). Best for startups, real-time Q&A, frequently changing data.
- nano-graphrag (2k+ stars): ~800 lines of code preserving entity extraction, community detection, and Local/Global queries. Swappable LLM/embedding modules, incremental insertion with md5 dedup, networkx + milvus-lite by default (Neo4j/FAISS/Ollama optional). Recomputes communities on every insert — learning tool, not production.
- KAG (8k+ stars, Ant Group + OpenKG): built on OpenSPG with the LLMFriSPG framework inspired by DIKW. Innovations: mutual indexing of knowledge and text chunks (auditable reasoning back to source evidence), a logical-form planner that decomposes questions into plan→reason→retrieve operator chains, and knowledge alignment via semantic reasoning for entity disambiguation. Unlike LLM-driven auto-extraction in GraphRAG, KAG uses schema constraints and structured, traceable reasoning. Native Chinese optimization. Best for healthcare, legal, finance where conclusions must be explainable.
- HippoRAG (3k+ stars, OSU NLP): LLM as "neocortex," knowledge graph + Personalized PageRank as "hippocampus." PPR probability diffusion from query entities enables multi-hop reasoning in a single retrieval step without iterative LLM calls. More academic; low engineering maturity.
- PathRAG (1k+ stars, BUPT GAMMA): a "denoising patch" for LightRAG. Streaming pruning: DFS to three-hop neighbors, then resource-allocation-based pruning of low-weight branches (nodes with more edges get more resources). Path-guided prompting describes graph paths in text ("X connects to Z via edge Y") instead of dumping graph data. Lower noise and token cost; an upgrade path for LightRAG users.
- Yuxi-Know (4k+ stars): application-layer platform on LightRAG with LangChain v1 + FastAPI + Vue full-stack GUI — dashboards, knowledge base and graph visualization, MinerU PDF parsing, user/department permissions, ScrapeGraphAI web scraping. For no-code internal knowledge bases.
- NebulaGraph / Neo4j RAG: infrastructure-layer graph databases rather than GraphRAG frameworks. NebulaGraph: distributed, storage-compute separated, trillion-scale edges, 99.999% availability. Neo4j: Cypher-based graph retrieval, first choice for relationship-dense apps.
- GraphRAG does not replace vector RAG. For finding a specific passage, classic vector retrieval is faster and more accurate. Hybrid retrieval often wins.
- Don't underestimate indexing costs. Per-chunk LLM extraction plus per-community summaries can cost an order of magnitude more than expected. Pilot on a small dataset first.
- Tune prompts for Chinese. Default prompts are English; entity extraction recall drops on Chinese text without tuning. KAG is natively optimized for Chinese.
- Don't put nano-graphrag in production. Community recomputation on every insert kills performance at scale.
- Watch LazyGraphRAG. Claimed 0.1% indexing cost could reshape the landscape if fully open-sourced into the main repo.
Comparison matrices (abridged)
| Dimension | Microsoft GraphRAG | LightRAG | nano-graphrag | KAG | HippoRAG | PathRAG | Yuxi-Know | |---|---|---|---|---|---|---|---| | Core algorithm | Leiden communities + summaries | Dual-level (node+edge) | Slim GraphRAG | Logical-form reasoning + DIKW | PPR diffusion | Streaming path pruning | Inherits LightRAG | | Incremental update | Weak (recomputes communities) | Strong (union merge) | Supported but recomputes | Strong | Moderate | Strong | Strong | | Multimodal | Moderate | Strong (RAG-Anything) | None | Moderate | None | None | Very strong (MinerU) | | Indexing cost | Very high | Low-mid | Mid-high | Mid | Mid | Low-mid | — | | Query token cost | Very high (Map-Reduce) | Low (two LLM rounds) | High | Mid | Low (single step) | Low | — | | Complex reasoning | Strong | Moderate | Strong | Very strong | Strong | Moderate | — | | Global insight | Strong | Moderate | Strong | Moderate | Weak | Moderate | — | | Chinese support | Prompt tuning needed | Prompt tuning needed | Prompt tuning needed | Native | Prompt tuning needed | Prompt tuning needed | Native | | Production readiness | High | High | Low | Mid | Low | Low | High |
Selection decision tree
1. Data scale: TB+/trillion edges → NebulaGraph + custom app layer. Millions to hundreds of millions → continue. Tens of thousands of docs → anything below works. 2. Query type: Global summaries → GraphRAG. Multi-hop relations → KAG (auditable) or HippoRAG (cheap). Specific entity lookups → LightRAG. Fast + accurate + cheap → LightRAG Hybrid or PathRAG. 3. Update frequency: Static → GraphRAG. Daily new data → LightRAG/PathRAG. Streaming → LightRAG. 4. Budget/team: Big budget + expert team → GraphRAG or KAG. Limited budget → LightRAG. No code → Yuxi-Know. Learning → nano-graphrag. 5. Domain: Auditable reasoning (medical/legal) → KAG. Global analysis (finance) → GraphRAG. Real-time Q&A → LightRAG. Multi-hop research → HippoRAG. Noise-sensitive → PathRAG.
A realistic hybrid architecture
1. Route ~90% of daily queries through LightRAG for speed and cost control. 2. Escalate complex analyses to GraphRAG when LightRAG's confidence is low or the user explicitly wants global analysis. 3. Share the underlying knowledge graph storage to avoid duplicate indexing.
Benefit: pay GraphRAG's costs only where truly needed. Downside: maintaining two systems requires dedicated staffing.
Common pitfalls
Trends
1. Collapsing costs: GraphRAG → LightRAG (1/12 cost) → LazyGraphRAG (three more orders of magnitude claimed). Adoption barrier keeps dropping. 2. Domain specialization: KAG's schema constraints + logical reasoning offer more ToB value than general-purpose frameworks. 3. Multimodal fusion: RAG-Anything (LightRAG) and MinerU (Yuxi-Know) extend "text to graph" toward "any format to graph." 4. Agentic GraphRAG: agents autonomously choosing between graph retrieval, vector retrieval, or both — LangGraph + LightRAG is already exploring this.
Appendix: quick links
| Project | GitHub | Paper | |---|---|---| | Microsoft GraphRAG | https://github.com/microsoft/graphrag | https://arxiv.org/abs/2404.16130 | | LightRAG | https://github.com/HKUDS/LightRAG | https://arxiv.org/abs/2410.05779 | | nano-graphrag | https://github.com/gusye1234/nano-graphrag | — | | KAG | https://github.com/OpenSPG/KAG | https://arxiv.org/abs/2409.13731 | | HippoRAG | https://github.com/OSU-NLP-Group/HippoRAG | https://arxiv.org/abs/2405.14831 | | PathRAG | https://github.com/BUPT-GAMMA/PathRAG | — | | Yuxi-Know | https://github.com/xerrors/Yuxi-Know | — | | NebulaGraph | https://github.com/vesoft-inc/nebula | — | | Neo4j RAG | https://github.com/neo4j/neo4j-graphrag-python | — |
> Data as of June 2026. Star counts and project status change constantly — refer to GitHub for real-time data.