Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · September 30, 2026 · repo

VectifyAI/PageIndexMafin 2.5ragagent-infrastructurelocal-ai

PageIndex Makes RAG Traceable, but the Trace You Get Depends on the Tier

VectifyAI's vectorless RAG repo gives page-level citations in open source and finer ones in its cloud, and its famous benchmark number belongs to a different product

Similarity is not relevance.

The PageIndex README makes that argument near the top ("similarity ≠ relevance"), and it's the whole pitch in four words. Vector RAG chops a document into chunks, embeds them, and hands the model whatever sits closest in embedding space. PageIndex throws that out. No vector database, no chunking. It builds a table-of-contents-style tree of the document and lets the model reason its way down the branches, the way you'd flip to "Item 7" in a 10-K instead of searching for words that sound like revenue.

VectifyAI/PageIndex was sitting at #7 on Trendshift's daily board this afternoon, with about 38k stars on a cache-busted shields badge and v0.2.20 tagged on September 28. Most coverage leads with one number: 98.7% on FinanceBench.

I think that number is the least interesting thing about the repo. And it isn't this repo's number.

The benchmark belongs to Mafin 2.5

Follow the README's FinanceBench link and you land in a separate repository, VectifyAI/Mafin2.5-FinanceBench. Its first line is clear: "Mafin2.5 is our latest RAG model on Financial reports, built on PageIndex."

So the 98.7% is Mafin 2.5's result. Mafin is Vectify's financial product. PageIndex is the retrieval framework underneath it. Those are different things, and when you pip install -U pageindex you get the framework, not the product that set the score.

The evaluation write-up has other details worth knowing. It says Mafin 2.5 covered the full benchmark, while several competitors in its table (Fintool, ChatGPT 4o with Search and Perplexity) were evaluated on about two-thirds of the questions, and it uses "expert human annotations" to handle questions it considered ambiguous or invalid. It also carries its own caveats: the benchmark "may contain inconsistencies, ambiguities, or errors in ground truth answers," and current evaluations emphasize "simple retrieval tasks based on a single document." I couldn't find a question count or a description of who graded the answers.

None of that makes 98.7% false. It makes it a vendor result on a vendor product, with the vendor choosing how to treat hard questions. File it that way.

The better reason to look: traceability

Here's what I'd actually care about. The README describes retrieval results as "traceable to explicit references." When the model walks a tree, the path it took is a record: this section, under this heading, on these pages. A vector top-k gives you five chunks and a similarity score. It can't tell you why those five, and the chunks often don't line up with anything a human would call a section.

That matters more this week than it did last month. Multi-source agents are getting called out for citing the wrong source for a true claim, and provenance is turning into something teams have to show, not just assert. A retriever that records the path it took hands you half the provenance record for free.

The mechanism has a nice split, too. According to the README, "the tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it." That's the PageIndex Flash method, now the default in local mode. The shape of your tree comes from the document's own layout. A model writes the summaries on each node. Then a chat model reasons over the tree at query time.

In the README's example, the index model defaults to gpt-5.6-luna and the chat model is gpt-5.6-sol, with an OPENAI_API_KEY set in the environment. So "local mode" means local indexing and storage. The reasoning still calls a hosted model unless you wire in something else, and the README I read only demonstrates OpenAI models.

Where the tiers split

This is the part the star count hides. PageIndex has a local mode, which is the open-source code in the repo, and a cloud mode, which needs a PageIndex API key. The README's comparison table draws the line like this:

Local mode (open source): text-based PDFs, local indexing and storage, page-level citations.

Cloud mode (paid): text-based, scanned and image-rich documents, managed indexing, OCR, image understanding, block-level citations, metadata, folders, and an MCP server.

Read that twice if you're planning an agent integration. The finer trace, block-level, lives in the cloud. Scanned PDFs, which describes a big share of real contracts, invoices and older filings, live in the cloud. And the MCP server, the thing that would let your agent call PageIndex as a tool, is listed as cloud-only.

That's a fair business model. MIT-licensed code, a paid service on top. It's also a real constraint, because the traceability pitch is strongest at block level, and the open-source version stops at the page.

A page-level citation on a 300-page filing still beats "chunk 4,812." It's still a page, though. If a page holds three tables and a footnote, your reviewer is reading all four to find the claim.

Who it fits, and who should skip it

PageIndex fits a specific shape of corpus: long, structured, text-based documents where the answer lives in one findable section. Annual reports, regulatory filings, standards, policy manuals, product documentation with a real hierarchy. Those documents already have a table of contents in their bones, and a tree built from layout will mirror it.

It fits badly on the opposite shape. Thousands of short documents (support tickets, chat logs, emails) don't have trees worth walking, and vector search across many small items is exactly what embeddings do well. Same for questions that are really "find every mention of X across the corpus," which is a search problem, not a navigation problem.

If your corpus is mixed, the honest answer is probably both: vector search to pick the document, tree reasoning to find the page inside it.

Put this into practice

The lowest-friction way to judge PageIndex is to run it against one document you already know well.

Start with a text-based PDF that has real structure. A 10-K, a standards document or a long technical manual with numbered sections. Install with pip install -U pageindex, set your OpenAI key, and build an index with the default local settings.

Look at the tree before you ask a question. Since the structure comes from layout, the tree tells you immediately whether PageIndex understood your document. If sections are missing or merged, retrieval will inherit that.

Ask five questions you already know the answers to, including one whose answer spans two sections. Check whether the page citations point where you'd look yourself.

Run the same five through your current vector pipeline. Score correctness, then time how long a reviewer takes to confirm each answer from the citation alone. That review time is the metric PageIndex is really competing on.

Count the model calls. Tree reasoning means the chat model reads summaries at each level. On a deep document that can be several calls per question. Watch your token bill for a day before you decide it's cheaper than embeddings.

If you need scans or an MCP tool, price the cloud tier up front. Don't build the agent on local mode and discover the gap at integration time.

Honest limitations

The benchmark caveat is the big one: the headline number is Mafin 2.5's, set on Vectify's own evaluation setup, not a reproduction of open-source PageIndex.

Layout-derived trees depend on layout. The README I read doesn't say how PageIndex handles documents without clear headings or structure, like a scanned letter, a transcript, or a slide export where every page is its own island. I'd expect weaker trees there, and on scans local mode doesn't run at all.

Local mode is not offline. The indexing summaries and the tree reasoning both call a model, and the documented path is OpenAI. If your reason for "local" is data residency, check exactly what leaves the machine during indexing.

The license has a small wrinkle: the LICENSE file names "Vectify AI" as the copyright holder, while the README footer says "PageIndex AI." It's MIT either way, but if your legal team tracks holders, flag it.

And this is a fast-moving 0.2.x project. Tree formats and defaults (the README says Flash is "now the default") can change between releases, so pin the version you test.

What to do with it

PageIndex is worth an afternoon for one reason, and it isn't 98.7%. It's that retrieval with a readable path is a better starting point for provenance than a pile of nearest neighbors. Whether the open-source tier gives you enough of that path, page-level on text PDFs, is a question only your documents can answer.

Pick the document your team argues about most. Index it. See whether the citations end the argument or start a new one.

Sources: VectifyAI/PageIndex README; VectifyAI/Mafin2.5-FinanceBench; Mafin 2.5 announcement; PageIndex docs; Trendshift.


Medium metadata

  • Title: PageIndex Makes RAG Traceable, but the Trace You Get Depends on the Tier
  • Subtitle: VectifyAI's vectorless RAG repo gives page-level citations in open source and finer ones in its cloud, and its famous benchmark number belongs to a different product
  • Tags: RAG, Retrieval Augmented Generation, Open Source, LLM, AI Agents
  • Canonical URL: fervorai.dev (import from the published article URL)
  • Reading time: about 7 minutes