PageIndex: Vectorless RAG Hits 98.7% on FinanceBench via Tree Reasoning

PageIndex replaces vector databases with hierarchical tree indexes and LLM reasoning for retrieval, achieving 98.7% accuracy on FinanceBench versus 50% for vector RAG, with indexing cost of $0.001 per page and traceable page-level citations.

Architecture Digest
Architecture Digest
Architecture Digest
PageIndex: Vectorless RAG Hits 98.7% on FinanceBench via Tree Reasoning

Project Introduction

PageIndex is a vectorless, chunk-free reasoning-based RAG engine. It replaces vector indexes with hierarchical tree indexes and lets an LLM reason along the tree to retrieve answers, yielding results traceable to explicit page- or block-level citations rather than opaque similarity scores.

A comparison table highlights the differences:

Indexing: Vector RAG uses vector indexes; PageIndex uses tree indexes.

Retrieval: Vector RAG relies on semantic similarity search; PageIndex uses LLM reasoning along the tree.

Results: Vector RAG produces opaque, "vibe" retrieval; PageIndex provides traceable citations.

Context: Vector RAG only sees the query vector; PageIndex can incorporate conversation history and domain knowledge.

Core Highlights

1. No Vectors, No Chunking — Tree Index Replaces Vector Index

PageIndex discards vector databases and chunking entirely. Retrieval depends on LLM reasoning, directly addressing the core flaw that semantic similarity ≠ relevance.

2. Two-Stage Retrieval — Traceable and Explainable

Index (build): Generates a tree structure per document from layout extraction, without LLM involvement; indexing uses only cheap base models.

Retrieve (query): An LLM agent reasons step-by-step down the tree to locate the correct section.

Results are traceable to explicit citations — page-level in local mode, block-level in cloud mode.

3. Cost Does Not Scale with Document Length

Local indexing costs approximately $0.001 per page ; a 1,000-page textbook indexes for about one dollar in minutes, and the index is reused for every query.

Indexing time grows linearly: official tests show 13 seconds for 9 pages up to 4.5 minutes for 1,098 pages.

Compared to stuffing the entire PDF into the model context, PageIndex is far cheaper: 52 pages cost 2.1× less, 420 pages cost 16.6× less, and 805 pages exceed the context window entirely — because PageIndex only reads the nodes its reasoning hits.

4. FinanceBench 98.7% SOTA, Publicly Reproducible

On the FinanceBench financial QA benchmark, PageIndex achieves 98.7% accuracy versus 50% for vector RAG — a 48-percentage-point gap. The team also released PageIndex-OSS-Benchmark for open reproduction; questions are designed so answers are explicitly stated in the text, eliminating "reasoning failure" as an excuse.

Core Principles and Architecture Differences

Vector RAG's Defect

Vector RAG approximates relevance with similarity. Semantic similarity works for short, general QA but fails on long, complex professional documents (financial reports, legal filings, regulatory submissions), producing false positives (similar but irrelevant) and false negatives (relevant but dissimilar).

PageIndex's Architecture Design (Inspired by AlphaGo)

Index layer: swap vectors for trees. Each document gets a tree index derived from layout; indexing cost is ultra-low ($0.001/page) and quality does not depend on strong models.

Retrieval layer: swap similarity for reasoning. An LLM agent reasons down the tree — analogous to how a human reads a report: check the table of contents, locate the chapter, then read the passage.

Result layer: swap probabilities for citations. Outputs are traceable to exact page/block references, meeting compliance and audit requirements for explainability.

The conclusion: the endgame for long-document retrieval is not better vectors but reasoning-based retrieval. PageIndex hands document reading back to a thinking LLM.

Quick Start

Environment: Pure Python, install with pip install -U pageindex.

Configure key: Local mode includes an LLM key ( OPENAI_API_KEY).

Launch client:

import os
from pageindex import PageIndexClient

os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(
    index="gpt-5.6-luna",   # tree-indexing model (base model suffices)
    chat="gpt-5.6-sol",     # tree-reasoning model (use strongest you can afford)
)
doc_id = client.submit_document("report.pdf")["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))

Dual-mode switch: Local mode handles text-based PDFs; scanned or image-heavy documents use Cloud — set index="cloud". Cloud hosts OCR, image understanding, tree indexing, storage, block-level citations, and MCP; the chat layer still uses your own model.

Team Deployment Solutions

1. Team Transformation Approach

Fits into the long-document QA slot: migrate retrieval of financial reports, legal docs, and regulatory filings from vector DB + chunking to tree index + reasoning retrieval. The integration boundary is clean — PageIndex only takes over the retrieval layer, leaving upstream document pipelines and downstream business apps untouched.

2. Deployment Options

Local SDK unifies the indexing pipeline: one tree per document, re-index on change (cost negligible at $0.001/page). Batch-build trees for existing corpora, incrementally index new docs; index assets persist in a local directory for team sharing.

3. Business System Integration

Plug PageIndex as a retrieval tool into existing agents. Official support for OpenAI Agents SDK, Claude Agent SDK, and MCP server enables one-line integration, giving agents tree-based long-document retrieval. Traceable citations feed directly into compliance and audit outputs.

4. Team Standards Customization

Unify index/chat model configs (index uses cheap base model, chat selects model per accuracy budget); choose citation granularity per compliance needs — page-level stays fully local, block-level + OCR requires Cloud.

Real Applicable Scenarios

Finance/Legal/Compliance teams: High-precision QA on financial reports, legal documents, regulatory filings with traceable results.

Research/Medical: Long-document retrieval over academic textbooks and medical literature; benchmark design (wrong answer = retrieval/read failure) targets these scenarios.

RAG teams avoiding vector DB maintenance: Eliminates vector DB, chunking, and embedding tuning entirely.

Audit scenarios needing explainable retrieval: Citations land on page/block, inherently explainable.

Not suitable for: Short-text or general knowledge QA (long-doc advantage unused); teams with no LLM key (retrieval quality depends on model strength).

Pros, Cons, and Pitfalls

Core Advantages

Novel yet practical approach (tree index + reasoning retrieval replaces similarity); hard data (FinanceBench 98.7% SOTA, public cost comparisons); cheap local indexing ($0.001/page); MIT licensed with active multi-product iteration.

Limitations

Local mode limited to text PDFs: Scanned or image-heavy docs require Cloud.

Quality hinges on model strength: Chat model must be the strongest affordable; weaker models visibly degrade retrieval.

Cloud features are paid: OCR, block-level citations, MCP, File System are commercial.

Ecosystem is new: Repository created 2025-04; audience is niche long-doc RAG; not for general QA.

Common Deployment Pitfalls

Confirm document type upfront — scanned PDFs go straight to Cloud, don't force local; don't downgrade chat model to save money, accuracy drops visibly; pin versions — ecosystem moves fast, today's working params may break tomorrow.

Author's Take

Vector RAG's fundamental flaw is conflating similarity with relevance. PageIndex discards the vector DB and uses tree indexes so an LLM can reason its way to the right page like a human — FinanceBench 98.7% vs. 50% proves this path works for long-document retrieval.

For teams doing long-doc QA, this is a serious new direction worth piloting; but local mode only ingests text PDFs, and retrieval quality is bound to the model you provide.

Open source: https://github.com/VectifyAI/PageIndex (MIT)

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

RAGOpen sourceLLM reasoningtree indexingPageIndexlong-document QAFinanceBenchvectorless
Architecture Digest
Written by

Architecture Digest

Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.