English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PageIndex Deep Dive: Vectorless RAG with Tree Search Hits 98.7% on FinanceBench

Forum topic · 小凯 · 2026-05-16

Summary

PageIndex, an open-source project by VectifyAI (MIT license, ~30k GitHub stars), replaces traditional vector-database RAG with a hierarchical tree index and LLM-driven reasoning-based navigation. Instead of chunking documents and matching embeddings by similarity, PageIndex parses PDFs, detects the table of contents, builds a hierarchical tree (with fallback modes for documents without a TOC), verifies the structure via a self-healing pipeline, and lets an LLM reason over node summaries to locate the exact pages containing the answer. The companion Mafin 2.5 financial RAG system reports 98.7% accuracy on the full FinanceBench benchmark, remaining constant whether powered by GPT-4o or DeepSeek v3 — suggesting the retrieval framework matters more than the underlying LLM. Key advantages: no vector database, no chunking, page-level citations, and fully auditable retrieval paths. Trade-offs include higher latency and token cost from multiple LLM calls. The post argues PageIndex is optimal for long, highly structured documents (finance, legal, academic), while vector RAG remains better for short, unstructured text — pointing toward hybrid routing architectures.

PageIndex Deep Dive: From Similarity Retrieval to Reasoning-Based Retrieval

> Project: PageIndex (GitHub: VectifyAI/PageIndex, ~30k stars) > Authors: Mingtian Zhang, Yu Tang, VectifyAI Team > License: MIT > Core claim: Similarity ≠ Relevance — retrieval needs reasoning, not vector similarity > Benchmark result: Mafin 2.5 financial RAG → 98.7% accuracy on FinanceBench

Key points

  • Core thesis: "Similarity ≠ Relevance. What we truly need in retrieval is relevance, and that requires reasoning." Vector search finds text that *looks like* the query; it does not find what *answers* the query.
  • Architecture: PageIndex builds a hierarchical tree index of documents (no vector DB, no chunking) and lets an LLM navigate the tree at query time, extracting only the target pages.
  • Pipeline: Parse → Detect TOC → Build Tree → Verify → Enrich → Search (PyMuPDF parsing, TOC detection, three tree-building modes, LLM-generated node summaries, tree-based search).
  • Self-healing verification: TOC accuracy is sampled and checked; below 60% accuracy it degrades to a simpler mode; between 60–100% it attempts repairs (up to 3 retries). This tolerates the notorious unreliability of PDF parsing (OCR errors, page-number offsets, broken tables).
  • Agentic retrieval flow: query → get document tree (structure + summaries only, saving tokens) → LLM reasons over branches → get_page_content(doc_id, "21-28") → answer grounded in exact pages.
  • Results: Mafin 2.5 reports 98.7% on the full 100% of FinanceBench (public results), vs. Quantly 94%, Fintool 98% (66.7% coverage), ChatGPT 4o + Search 31%, Perplexity 45%. Accuracy holds at 98.7% with both GPT-4o and DeepSeek v3 backends — implying the retrieval framework contributes more than the underlying LLM.
  • Caveats: FinanceBench question design is not fully public, so independent verification is limited; ChatGPT's low score reflects its retrieval architecture (web/vector search), not LLM capability.
  • Vector RAG vs. PageIndex

    | Dimension | Vector RAG | PageIndex | |---|---|---| | Retrieval | Vector similarity matching | LLM reasoning over tree navigation | | Document handling | Chunking (manual/semantic) | Natural hierarchical structure preserved | | Storage | Vector DB (Pinecone/Weaviate/Chroma) | None (plain JSON index) | | Explainability | Low ("this chunk is most similar") | High ("choose Risk Factors → p21-28") | | Latency | Milliseconds | Seconds (multiple LLM calls) | | Token cost | Near zero at retrieval | LLM calls for indexing and search | | Page citations | Usually not possible | Exact page/section references | | Auditability | Weak | Strong (traceable retrieval path) |

    Best-fit scenarios: PageIndex excels on long, structured documents (10-K/10-Q filings, legal contracts, academic papers). Vector RAG remains superior for short or unstructured text (chat logs, code, general knowledge bases). The likely end state is a hybrid architecture that routes queries to the appropriate retriever.

    The AlphaGo analogy

    PageIndex cites AlphaGo as inspiration: LLM judgment of promising branches plays the role of a policy network, node summaries act as value estimates, and hierarchical navigation mirrors selective tree search. Unlike AlphaGo's learned intuition, however, PageIndex's search is explicit, symbolic, and interpretable — good for auditability, but every step costs an LLM call.

    Limitations and outlook

  • Latency: seconds per query; mitigations include tree caching and precomputed node embeddings for pruning.
  • Cost: indexing and retrieval both require LLM calls; small models could prune, large models decide.
  • TOC-less documents: auto-generated trees have uncertain quality.
  • PDF parsing: PyMuPDF-based; complex layouts remain error-prone — vision models could help.
  • Scope: unsuited to code, chat logs; cross-document queries require extending the single-document tree into a document relationship graph.
  • Conclusion

    PageIndex's value is less the 98.7% number than the paradigm shift: a document is a tree, not a pile of chunks; retrieval is reasoning-based navigation, not similarity ranking. It is not a death sentence for vector databases — it is a reminder that structured knowledge needs structured retrieval.

    References

  • PageIndex GitHub: https://github.com/VectifyAI/PageIndex
  • Mafin 2.5 FinanceBench results: https://github.com/VectifyAI/Mafin2.5-FinanceBench
  • TypeScript port: https://github.com/FutureSpeakAI/agent-fridays-pageindex-rag
  • VectifyAI: https://vectify.ai
  • FinanceBench paper: https://arxiv.org/abs/2311.11944 (Patel et al., 2023)

Tags

#rag#pageindex#vectorless-rag#tree-search#financebench#llm#document-understanding#retrieval-augmented-generation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620127