PageIndex Deep Dive: From Similarity Retrieval to Reasoning-Based Retrieval
> Project: PageIndex (GitHub: VectifyAI/PageIndex, ~30k stars) > Authors: Mingtian Zhang, Yu Tang, VectifyAI Team > License: MIT > Core claim: Similarity ≠ Relevance — retrieval needs reasoning, not vector similarity > Benchmark result: Mafin 2.5 financial RAG → 98.7% accuracy on FinanceBench
Key points
- Core thesis: "Similarity ≠ Relevance. What we truly need in retrieval is relevance, and that requires reasoning." Vector search finds text that *looks like* the query; it does not find what *answers* the query.
- Architecture: PageIndex builds a hierarchical tree index of documents (no vector DB, no chunking) and lets an LLM navigate the tree at query time, extracting only the target pages.
- Pipeline:
Parse → Detect TOC → Build Tree → Verify → Enrich → Search(PyMuPDF parsing, TOC detection, three tree-building modes, LLM-generated node summaries, tree-based search). - Self-healing verification: TOC accuracy is sampled and checked; below 60% accuracy it degrades to a simpler mode; between 60–100% it attempts repairs (up to 3 retries). This tolerates the notorious unreliability of PDF parsing (OCR errors, page-number offsets, broken tables).
- Agentic retrieval flow: query → get document tree (structure + summaries only, saving tokens) → LLM reasons over branches →
get_page_content(doc_id, "21-28")→ answer grounded in exact pages. - Results: Mafin 2.5 reports 98.7% on the full 100% of FinanceBench (public results), vs. Quantly 94%, Fintool 98% (66.7% coverage), ChatGPT 4o + Search 31%, Perplexity 45%. Accuracy holds at 98.7% with both GPT-4o and DeepSeek v3 backends — implying the retrieval framework contributes more than the underlying LLM.
- Caveats: FinanceBench question design is not fully public, so independent verification is limited; ChatGPT's low score reflects its retrieval architecture (web/vector search), not LLM capability.
- Latency: seconds per query; mitigations include tree caching and precomputed node embeddings for pruning.
- Cost: indexing and retrieval both require LLM calls; small models could prune, large models decide.
- TOC-less documents: auto-generated trees have uncertain quality.
- PDF parsing: PyMuPDF-based; complex layouts remain error-prone — vision models could help.
- Scope: unsuited to code, chat logs; cross-document queries require extending the single-document tree into a document relationship graph.
- PageIndex GitHub: https://github.com/VectifyAI/PageIndex
- Mafin 2.5 FinanceBench results: https://github.com/VectifyAI/Mafin2.5-FinanceBench
- TypeScript port: https://github.com/FutureSpeakAI/agent-fridays-pageindex-rag
- VectifyAI: https://vectify.ai
- FinanceBench paper: https://arxiv.org/abs/2311.11944 (Patel et al., 2023)
Vector RAG vs. PageIndex
| Dimension | Vector RAG | PageIndex | |---|---|---| | Retrieval | Vector similarity matching | LLM reasoning over tree navigation | | Document handling | Chunking (manual/semantic) | Natural hierarchical structure preserved | | Storage | Vector DB (Pinecone/Weaviate/Chroma) | None (plain JSON index) | | Explainability | Low ("this chunk is most similar") | High ("choose Risk Factors → p21-28") | | Latency | Milliseconds | Seconds (multiple LLM calls) | | Token cost | Near zero at retrieval | LLM calls for indexing and search | | Page citations | Usually not possible | Exact page/section references | | Auditability | Weak | Strong (traceable retrieval path) |
Best-fit scenarios: PageIndex excels on long, structured documents (10-K/10-Q filings, legal contracts, academic papers). Vector RAG remains superior for short or unstructured text (chat logs, code, general knowledge bases). The likely end state is a hybrid architecture that routes queries to the appropriate retriever.
The AlphaGo analogy
PageIndex cites AlphaGo as inspiration: LLM judgment of promising branches plays the role of a policy network, node summaries act as value estimates, and hierarchical navigation mirrors selective tree search. Unlike AlphaGo's learned intuition, however, PageIndex's search is explicit, symbolic, and interpretable — good for auditability, but every step costs an LLM call.
Limitations and outlook
Conclusion
PageIndex's value is less the 98.7% number than the paradigm shift: a document is a tree, not a pile of chunks; retrieval is reasoning-based navigation, not similarity ranking. It is not a death sentence for vector databases — it is a reminder that structured knowledge needs structured retrieval.