English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RAG-Anything: Extending LightRAG to All-Modal Retrieval with Dual Knowledge Graphs

Forum topic · 小凯 · 2026-05-23

Summary

RAG-Anything, from HKU's HKUDS lab (arXiv:2510.12323), extends LightRAG's graph-based retrieval to full multimodal documents containing images, tables, and equations. The system first decomposes documents into atomic content units with contextual anchoring, then builds two graphs: a cross-modal knowledge graph, where a multimodal LLM generates detailed descriptions and entity summaries for non-text units, and a text knowledge graph built via standard entity-relation extraction. The two graphs are merged through entity alignment. Retrieval combines structural graph navigation with dense semantic matching, and a VLM generates the final answer conditioned on both text context and recovered visual content. On DocBench (63.4% vs LightRAG 58.4%) and MMLongBench (42.8% vs 38.9%), RAG-Anything outperforms baselines, with its advantage widening on documents over 100 pages. Ablations show graph construction is essential, while reranking adds modest gains. Acknowledged limitations include text-centric retrieval bias and difficulties with complex layouts, plus open questions on MLLM description reliability, entity alignment robustness, and indexing cost.

RAG-Anything: When LightRAG Opens Its "All-Modal" Eyes

> Repository: https://github.com/HKUDS/RAG-Anything > Paper: arXiv:2510.12323v1 > Authors: Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang (University of Hong Kong)

1. The Problem: Why Text-Only RAG Falls Short

Existing RAG frameworks assume knowledge bases are pure text. That assumption collapses in the real world: academic papers contain figures and formulas, financial reports contain tables and trend charts, medical records contain imaging and diagnostic metrics. Forcing everything into text loses as much information as describing a painting to someone who has never seen color.

HKU's HKUDS team previously released LightRAG, which pushed text RAG efficiency to a new level with dual-level retrieval (keywords + semantics). But LightRAG cannot handle images, tables, or equations — an architectural blind spot, not just a missing feature.

RAG-Anything was born from this insight. Rather than starting from scratch, it extends LightRAG's graph retrieval ideas to all modalities.

2. Core Design: Dual-Graph Construction and Hybrid Retrieval

2.1 Atomic Content Units

The first step decomposes documents into "atomic content units":

\[c_j = (t_j, x_j)\]

where \(t_j\) is the modality type (text/image/table/equation) and \(x_j\) is the extracted raw content. MinerU handles high-fidelity extraction — multi-column layouts, nested tables, and inter-line formulas in PDFs are all structured.

Key design: contextual anchoring. Figures stay connected to captions, equations to surrounding definitions, tables to explanatory text. These anchors later become edges in the graph.

2.2 Dual-Graph: Not One Graph, But Two Merged

Cross-Modal Knowledge Graph — For non-text units (figures, tables, equations), an MLLM generates two textual representations:

  • Detailed description \(d_j^{chunk}\): used for retrieval matching
  • Entity summary \(e_j^{entity}\): used for graph construction
  • Generation is context-aware: each unit is processed with its local neighborhood \(C_j = \{c_k \mid |k-j| \leq \delta\}\), ensuring descriptions reflect the unit's role in the document structure. A graph extraction routine \(R(\cdot)\) then extracts fine-grained entities and relations:

    \[(\mathcal{V}_j, \mathcal{E}_j) = R(d_j^{chunk})\]

    Each non-text unit gets a multimodal entity node \(v_j^{mm}\), anchoring intra-chunk entities via explicit belongs_to edges.

    Text Knowledge Graph — For text units, the LightRAG/GraphRAG pipeline is used directly: NER + relation extraction.

    Graph Fusion — The two graphs merge via entity alignment, using entity names as matching keys to unify semantically equivalent entities into a single knowledge graph \(\mathcal{G} = (\mathcal{V}, \mathcal{E})\).

    2.3 Hybrid Retrieval: Structure Navigation + Semantic Matching

    (i) Structural knowledge navigation — Keywords from the query are matched exactly in \(\mathcal{G}\), then neighborhoods are strategically expanded within a hop limit. This captures cross-modal multi-hop reasoning paths (e.g., from a paragraph to the figure it cites, to a specific panel).

    (ii) Semantic similarity matching — Simultaneously, dense vector similarity search runs between the query embedding \(\mathbf{e}_q\) and all embedded components (atomic blocks, graph entities, relation representations).

    The candidate pools merge as \(\mathcal{C}(q) = \mathcal{C}_{stru}(q) \cup \mathcal{C}_{seman}(q)\), ranked by multi-signal fusion scoring (structural importance + semantic similarity + query-inferred modality preference).

    2.4 Synthesis: VLM Conditional Generation

    Retrieved candidates are split:

  • Text components: entity summaries, relation descriptions, chunk content assembled into \(\mathcal{P}(q)\)
  • Visual components: original visual content recovered by dereferencing into \(\mathcal{V}^*(q)\)
  • The VLM generates jointly conditioned on query, text context, and visual content:

    \[\text{Response} = \text{VLM}(q, \mathcal{P}(q), \mathcal{V}^*(q))\]

    3. Experiments: Overwhelming Advantage on Long Documents

    | Dataset | Docs | Avg. Pages | Avg. Tokens | Questions | |---------|------|-----------|-------------|-----------| | DocBench | 229 | 66 | 46,377 | 1,102 | | MMLongBench | 135 | 47.5 | 21,214 | 1,082 |

    Overall accuracy, DocBench: GPT-4o-mini 51.2%, LightRAG 58.4%, MMGraphRAG 61.0%, RAG-Anything 63.4%

    Overall accuracy, MMLongBench: GPT-4o-mini 33.5%, LightRAG 38.9%, MMGraphRAG 37.7%, RAG-Anything 42.8%

    Long-document effect — the advantage grows with document length:

  • DocBench >100 pages: 68.2% vs MMGraphRAG 54.6% (+13.6 pts)
  • DocBench >200 pages: 68.8% vs 55.0% (+13.8 pts)
  • MMLongBench 51–100 pages: +9.3 pts; 101–200 pages: +7.9 pts
  • This validates the core hypothesis: dual-graph construction and cross-modal hybrid retrieval are especially effective for scattered multimodal evidence in long documents.

    Ablations (DocBench): Chunk-only (no graph) 60.0%, w/o reranker 62.4%, full system 63.4%. Graph construction is essential (-3.4 pts without it); reranking helps modestly, meaning the core gains come from graph retrieval and cross-modal integration.

    4. Case Studies

    Multi-panel figure interpretation: For t-SNE visualizations with multiple subpanels, the constructed visual layout graph includes panels, axis titles, legends, and captions as nodes, with edges like panel_contains_plot, caption_provides_context, and subfigure_relates_hierarchically. This guides retrieval to the correct panel.

    Financial table navigation: For finding the intersection of the "Wages and salaries" row and "2020" column (26,778 million), the table becomes a graph of row headers, column headers, data cells, and units, with row-of, column-of, header-applies-to, and unit-of edges. Baselines treating tables as linear text often confuse numeric ranges and years.

    5. Relation to Prior Work

  • LightRAG: RAG-Anything is the modal extension of LightRAG; it inherits text graph construction and dual-level retrieval, and outperforms LightRAG as a baseline.
  • MMGraphRAG: only handles basic images; tables and equations are treated as plain text, losing structural information.
  • VisRAG: preserves document layout as images but lacks fine-grained relation modeling.
  • VideoRAG: handles the video modality; complementary to RAG-Anything's document focus.
  • 6. Limitations and Future Directions

    The paper's appendix candidly notes two fundamental issues:

    1. Text-centric retrieval bias: even for queries explicitly asking for visual information, the system prefers text sources — a fundamental weakness of cross-modal attention. 2. Document structure challenges: complex layouts and non-linear information flow remain hard; merged cells and ambiguous table boundaries still cause errors.

    7. Engineering Perspective

    Dependencies: MinerU for parsing (PDF/images/DOCX/PPTX/XLSX), MLLM for description generation (GPT-4o-mini in the paper), text-embedding-3-large (3072-dim) embeddings, bge-reranker-v2-m3 reranker, and a VLM for final generation.

    Key configurations: entity/relation token limit 20,000, chunk token limit 12,000; modality-aware query encoding detecting keywords like "figure/chart/table/equation"; fully offline document parsing and graph construction.

    Extensibility: custom parser plugins (v1.2.10+), multilingual prompt templates, local model backends via Ollama and vLLM, compatibility with LangChain and LlamaIndex.

    8. Open Questions

    1. MLLM description reliability: hallucinated descriptions would propagate through graph construction and retrieval; the paper does not quantify this. 2. Entity alignment precision: name-based matching may be fragile with abbreviations, synonyms, and cross-lingual cases. 3. Compute cost: two MLLM calls per non-text unit, dual-graph construction, and cross-modal reranking — will indexing cost become a deployment bottleneck for long documents? 4. Modality extensibility: how well do these representations generalize to new modalities (3D models, code blocks, interactive charts)? 5. Competition with end-to-end VLMs: as VLM context windows grow (e.g., Gemini 1M+ tokens), will the advantage of explicit graph structure shrink?

    References

  • Guo et al. (2025). RAG-Anything: All-in-One RAG Framework. arXiv:2510.12323.
  • Guo et al. (2024). LightRAG: Simple and Fast Retrieval-Augmented Generation. arXiv:2410.05779.
  • Wang et al. (2024). MinerU: An Open-Source Solution for Precise Document Content Extraction. arXiv:2409.18839.
  • Zou et al. (2024). DocBench: A Benchmark for Evaluating LLM-based Document Reading Systems. arXiv:2407.10701.
  • Ma et al. (2024). MMLongBench-Doc: Benchmarking Long-Context Document Understanding with Visualizations. NeurIPS 37.

Tags

#rag#multimodal-ai#lightrag#knowledge-graph#document-understanding#retrieval-augmented-generation#hku#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620710