RAG-Anything: When LightRAG Opens Its "All-Modal" Eyes
> Repository: https://github.com/HKUDS/RAG-Anything > Paper: arXiv:2510.12323v1 > Authors: Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang (University of Hong Kong)
1. The Problem: Why Text-Only RAG Falls Short
Existing RAG frameworks assume knowledge bases are pure text. That assumption collapses in the real world: academic papers contain figures and formulas, financial reports contain tables and trend charts, medical records contain imaging and diagnostic metrics. Forcing everything into text loses as much information as describing a painting to someone who has never seen color.
HKU's HKUDS team previously released LightRAG, which pushed text RAG efficiency to a new level with dual-level retrieval (keywords + semantics). But LightRAG cannot handle images, tables, or equations — an architectural blind spot, not just a missing feature.
RAG-Anything was born from this insight. Rather than starting from scratch, it extends LightRAG's graph retrieval ideas to all modalities.
2. Core Design: Dual-Graph Construction and Hybrid Retrieval
2.1 Atomic Content Units
The first step decomposes documents into "atomic content units":
where \(t_j\) is the modality type (text/image/table/equation) and \(x_j\) is the extracted raw content. MinerU handles high-fidelity extraction — multi-column layouts, nested tables, and inter-line formulas in PDFs are all structured.
Key design: contextual anchoring. Figures stay connected to captions, equations to surrounding definitions, tables to explanatory text. These anchors later become edges in the graph.
2.2 Dual-Graph: Not One Graph, But Two Merged
Cross-Modal Knowledge Graph — For non-text units (figures, tables, equations), an MLLM generates two textual representations:
- Detailed description \(d_j^{chunk}\): used for retrieval matching
- Entity summary \(e_j^{entity}\): used for graph construction
- Text components: entity summaries, relation descriptions, chunk content assembled into \(\mathcal{P}(q)\)
- Visual components: original visual content recovered by dereferencing into \(\mathcal{V}^*(q)\)
- DocBench >100 pages: 68.2% vs MMGraphRAG 54.6% (+13.6 pts)
- DocBench >200 pages: 68.8% vs 55.0% (+13.8 pts)
- MMLongBench 51–100 pages: +9.3 pts; 101–200 pages: +7.9 pts
- LightRAG: RAG-Anything is the modal extension of LightRAG; it inherits text graph construction and dual-level retrieval, and outperforms LightRAG as a baseline.
- MMGraphRAG: only handles basic images; tables and equations are treated as plain text, losing structural information.
- VisRAG: preserves document layout as images but lacks fine-grained relation modeling.
- VideoRAG: handles the video modality; complementary to RAG-Anything's document focus.
- Guo et al. (2025). RAG-Anything: All-in-One RAG Framework. arXiv:2510.12323.
- Guo et al. (2024). LightRAG: Simple and Fast Retrieval-Augmented Generation. arXiv:2410.05779.
- Wang et al. (2024). MinerU: An Open-Source Solution for Precise Document Content Extraction. arXiv:2409.18839.
- Zou et al. (2024). DocBench: A Benchmark for Evaluating LLM-based Document Reading Systems. arXiv:2407.10701.
- Ma et al. (2024). MMLongBench-Doc: Benchmarking Long-Context Document Understanding with Visualizations. NeurIPS 37.
Generation is context-aware: each unit is processed with its local neighborhood \(C_j = \{c_k \mid |k-j| \leq \delta\}\), ensuring descriptions reflect the unit's role in the document structure. A graph extraction routine \(R(\cdot)\) then extracts fine-grained entities and relations:
Each non-text unit gets a multimodal entity node \(v_j^{mm}\), anchoring intra-chunk entities via explicit belongs_to edges.
Text Knowledge Graph — For text units, the LightRAG/GraphRAG pipeline is used directly: NER + relation extraction.
Graph Fusion — The two graphs merge via entity alignment, using entity names as matching keys to unify semantically equivalent entities into a single knowledge graph \(\mathcal{G} = (\mathcal{V}, \mathcal{E})\).
2.3 Hybrid Retrieval: Structure Navigation + Semantic Matching
(i) Structural knowledge navigation — Keywords from the query are matched exactly in \(\mathcal{G}\), then neighborhoods are strategically expanded within a hop limit. This captures cross-modal multi-hop reasoning paths (e.g., from a paragraph to the figure it cites, to a specific panel).
(ii) Semantic similarity matching — Simultaneously, dense vector similarity search runs between the query embedding \(\mathbf{e}_q\) and all embedded components (atomic blocks, graph entities, relation representations).
The candidate pools merge as \(\mathcal{C}(q) = \mathcal{C}_{stru}(q) \cup \mathcal{C}_{seman}(q)\), ranked by multi-signal fusion scoring (structural importance + semantic similarity + query-inferred modality preference).
2.4 Synthesis: VLM Conditional Generation
Retrieved candidates are split:
The VLM generates jointly conditioned on query, text context, and visual content:
3. Experiments: Overwhelming Advantage on Long Documents
| Dataset | Docs | Avg. Pages | Avg. Tokens | Questions | |---------|------|-----------|-------------|-----------| | DocBench | 229 | 66 | 46,377 | 1,102 | | MMLongBench | 135 | 47.5 | 21,214 | 1,082 |
Overall accuracy, DocBench: GPT-4o-mini 51.2%, LightRAG 58.4%, MMGraphRAG 61.0%, RAG-Anything 63.4%
Overall accuracy, MMLongBench: GPT-4o-mini 33.5%, LightRAG 38.9%, MMGraphRAG 37.7%, RAG-Anything 42.8%
Long-document effect — the advantage grows with document length:
This validates the core hypothesis: dual-graph construction and cross-modal hybrid retrieval are especially effective for scattered multimodal evidence in long documents.
Ablations (DocBench): Chunk-only (no graph) 60.0%, w/o reranker 62.4%, full system 63.4%. Graph construction is essential (-3.4 pts without it); reranking helps modestly, meaning the core gains come from graph retrieval and cross-modal integration.
4. Case Studies
Multi-panel figure interpretation: For t-SNE visualizations with multiple subpanels, the constructed visual layout graph includes panels, axis titles, legends, and captions as nodes, with edges like panel_contains_plot, caption_provides_context, and subfigure_relates_hierarchically. This guides retrieval to the correct panel.
Financial table navigation: For finding the intersection of the "Wages and salaries" row and "2020" column (26,778 million), the table becomes a graph of row headers, column headers, data cells, and units, with row-of, column-of, header-applies-to, and unit-of edges. Baselines treating tables as linear text often confuse numeric ranges and years.
5. Relation to Prior Work
6. Limitations and Future Directions
The paper's appendix candidly notes two fundamental issues:
1. Text-centric retrieval bias: even for queries explicitly asking for visual information, the system prefers text sources — a fundamental weakness of cross-modal attention. 2. Document structure challenges: complex layouts and non-linear information flow remain hard; merged cells and ambiguous table boundaries still cause errors.
7. Engineering Perspective
Dependencies: MinerU for parsing (PDF/images/DOCX/PPTX/XLSX), MLLM for description generation (GPT-4o-mini in the paper), text-embedding-3-large (3072-dim) embeddings, bge-reranker-v2-m3 reranker, and a VLM for final generation.
Key configurations: entity/relation token limit 20,000, chunk token limit 12,000; modality-aware query encoding detecting keywords like "figure/chart/table/equation"; fully offline document parsing and graph construction.
Extensibility: custom parser plugins (v1.2.10+), multilingual prompt templates, local model backends via Ollama and vLLM, compatibility with LangChain and LlamaIndex.
8. Open Questions
1. MLLM description reliability: hallucinated descriptions would propagate through graph construction and retrieval; the paper does not quantify this. 2. Entity alignment precision: name-based matching may be fragile with abbreviations, synonyms, and cross-lingual cases. 3. Compute cost: two MLLM calls per non-text unit, dual-graph construction, and cross-modal reranking — will indexing cost become a deployment bottleneck for long documents? 4. Modality extensibility: how well do these representations generalize to new modalities (3D models, code blocks, interactive charts)? 5. Competition with end-to-end VLMs: as VLM context windows grow (e.g., Gemini 1M+ tokens), will the advantage of explicit graph structure shrink?