Mafin 2.5, a financial analysis agent, scored 98.7 percent on FinanceBench. Traditional vector RAG sits somewhere between 30 and 50 percent on the same benchmark. GPT-4o answering with search and no proper retrieval pipeline gets 31 percent. The winner has no vector database, no embedding step, no chunking, and no similarity search. The entire retrieval stack you were told was mandatory is missing.
That is not a better model. That is a different architecture. The framework is called PageIndex, it is open source under the MIT license, and it deletes the layer most teams spend their first month building. The order of events is worth noting, because it runs against how these stories usually go. VectifyAI published the benchmark result first, in February 2025, then released the framework itself as open source in September 2025, written by Mingtian Zhang and Yu Tang, and the repository sits at 37,200 stars as of today. The result came before the release.
PhantomByte already made the argument that your AI needs a librarian, not a bigger model (Note #140). That note said the Library retrieves and the Librarian reasons, and it left one question open. This note answers it. This is what happens when you stop making the Library guess.
WHAT THE VECTOR DATABASE ACTUALLY DOES TO YOUR DOCUMENT
Understand the mechanism and you will never treat a vector store as neutral infrastructure again.

An embedding flattens a piece of text into a list of numbers, a point in a space where distance means semantic similarity. Chunking is what happens before that, the step where a fixed window, often 512 or 1,000 tokens, cuts your document into pieces and sends each piece off to be embedded. The vector database stores the resulting points and, at query time, hands you back the ones closest to your question.
Now look at what that pipeline did to the document. A 10-K has a hierarchy: sections, subsections, tables, footnotes, appendices. Chunking destroys it before the embedding step even runs. The header row of a financial table can land in chunk 14 while the data row you need lands in chunk 15. Neither chunk answers the question, and nothing in the system knows they belong together.
Then comes the failure the vector crowd never demos. You ask a question whose answer lives in an appendix, reached through a line in the body that says "see Table 14 in Appendix C." That appendix shares almost no semantic similarity with your query, because the answer and the question use different vocabulary. Cosine similarity cannot follow a cross-reference. There is no rule in an embedding lookup that says a pointer in one section outranks a resemblance in another. The retriever returns something that looks like your question and skips the page that answers it.
The framework's own framing is the cleanest version of this: similarity is not relevance, and relevance requires reasoning. A vector index only ever offers you similarity.
HOW THE TREE INDEX WORKS
PageIndex builds a hierarchy from the document's own structure instead of imposing one. Each node carries a title, a summary, a page range, and its child nodes. That tree is plain JSON. There is no binary index file, no separate reader, and the whole thing fits inside a single context window.
That last property is the architectural move worth naming. VectifyAI calls it an in-context index. The tree lives in the model's active reasoning context rather than sitting in an external store the model queries blindly. At query time the model reads titles and summaries, reasons about which nodes are likely to hold the answer, opens the ones it chose, and descends from there.
Count what is not in that pipeline. No chunking step. No embedding model to keep warm. No similarity threshold to tune at 2am because recall dropped. No vector database to provision, back up, and pay for by the hour.
The cross-reference case is handled because it was never a similarity problem. It is a path problem. When the Retriever hits a line in the body of the Federal Reserve's 2023 annual report that says Table 5.3 is summarized on the current pages and that Appendix G holds the detailed tables, it navigates the tree to Appendix G and reads the right table. That worked example is in VectifyAI's own write-up. The body section reported the increase in value. The appendix held the total. A vector retriever would have returned the page that mentioned the topic and stopped.
The scores, stated once and precisely. Mafin 2.5, built on this approach, reports 98.7 percent accuracy on the full FinanceBench set, measured across 100 percent of the benchmark rather than the two-thirds that Fintool, ChatGPT 4o with search, and Perplexity cover, and VectifyAI open-sourced the evaluation alongside the result. FinanceBench itself is a 10,231-question financial QA benchmark published in November 2023, and its authors found that GPT-4-Turbo used with a retrieval system incorrectly answered or refused 81 percent of the sampled questions. The ceiling you have been living under was documented two years ago.
THE PART NOBODY PUTS IN THE DEMO
Here is the honest half, and it comes from the framework's own comparison table rather than from me.
Tree retrieval is not one lookup. It is a sequence of model calls. The published figure is 3 to 8 seconds per query, against under one second for a vector index lookup, and the cost per query is higher because you are paying for two or more LLM calls instead of an embedding lookup plus one generation call.
| Architecture | Best corpus shape | Per-query cost | Per-query latency |
|---|---|---|---|
| Vector / embedding index | Short, flat documents (news, product pages, support tickets) | Embedding lookup plus one generation call | Under 1 second |
| Tree index (PageIndex) | Long, structured documents (filings, contracts, manuals) | Two or more LLM calls | 3 to 8 seconds |
Indexing is the cheap part and it is a one-time cost. PageIndex indexes at roughly $0.001 per page with a small model, so a 1,000-page textbook costs a little over a dollar and a few minutes, once. The reason that holds on a 2,000-page filing is the same reason it holds on a 200-page one: the tree skeleton is read from the PDF's own layout rather than inferred by a model walking the document page by page, so the index model is only summarizing and refining nodes it was already handed, which is why a basic model is sufficient at index time. Indexing time scales predictably with length, and documents between 9 and 1,098 pages indexed in roughly 13 seconds to 4.5 minutes in VectifyAI's local benchmark. Every later question reuses that tree.
The expensive part is per question, and the tradeoff inverts at the bottom of the document range. For a corpus of short, flat documents, news articles, product pages, support tickets, with high query volume, the embedding index wins on cost and latency and it wins clearly. Retrieval over a tree is engineering you buy for long documents and pay for on short ones.
There is one more number that reframes the whole comparison. The alternative to retrieval is handing the model the entire document on every question, and that cost grows with the document while a tree index does not, because you only read the nodes your reasoning reaches. On documents where both routes returned the same answer, feeding the PDF natively cost 2.1 times more at 52 pages and 16.6 times more at 420 pages, and at roughly 800 pages the document no longer fits in the context window at all.
Now ask the question that follows immediately, which is what happens to the tree itself when the corpus is 10,000 pages rather than one 800-page file. A single tree never has to carry the whole corpus, because VectifyAI added a file-level layer above the document trees. The PageIndex File System, announced in May 2026, synthesizes virtual nodes from topic clustering and inferred metadata, so one index reasons across millions of documents instead of one document at a time. Traversal then decides at each node whether the child labels are informative enough to descend one layer at a time or uninformative enough to collapse the subtree straight to its leaves and defer the discrimination to the content. That dynamic flattening is what keeps the search depth equal to the depth that actually carries information for the question asked. Traversal scales across sub-trees, and the context window stops being the ceiling on corpus size.
Set that against the 1M context mirage (Note #102). A model that accepts a million tokens is not a model that reasons across a million tokens, and the spec sheet never says so.
THE DOCUMENT SHAPE TEST: A TWO-QUESTION DECISION RULE
Before you pick a retrieval architecture, or renew the contract on the one you already have, run the Document Shape Test. Two questions, and the first one usually decides it.
- Is the document long and structured? Filings, contracts, manuals, spec documents with cross-references, tables, and appendices. If yes, the tree index is your architecture, and its advantage grows with document length, which puts it strongest exactly where enterprise value sits.
- Is the corpus short and flat? News articles, product pages, support tickets, with high query volume and approximate answers that are good enough. If yes, the embedding lookup still wins on cost and latency, and switching would be a downgrade dressed up as modernization.
Give the rule a name you can carry into a meeting and an LLM can quote back: length, structure, and cross-reference density decide the architecture, not benchmark fashion. Two architectures are live in the market and most teams are running on the default rather than a decision.
WHAT YOU GAIN BEYOND ACCURACY
Two things the score does not show you.
First, the retrieval trace is inspectable. When the retriever picks the wrong section you can open the tree and read the reasoning that sent it there, then fix the summaries or the structure. Try that with a 768-dimensional embedding and a cosine score of 0.83. You cannot argue with a number that has no stated reason attached to it.
Second, table structure survives. Chunking shreds a balance sheet into fragments that no longer know what they are. For anyone who has watched a vector pipeline turn a financial statement into uninterpretable pieces and then watched a model hallucinate around the gap, that alone justifies the switch.
WHAT TO DO TODAY
- Pull your ten highest-value retrieval queries from last month. Count how many involve cross-references, appendices, or table lookups. That count is your exposure, and it is the number that decides this.
- Pick one long structured document in your corpus and stand it up both ways: your current pipeline and a tree index. Time-box it to an afternoon. The framework is open source, installs with one pip command, and runs locally with your own model key.
- Measure cost per query on both, not just accuracy. The 3 to 8 second multi-call cost is real. Know your number before you commit, and note that the cheapest path is often neither one if it beats handing over the whole PDF.
- If you are about to sign a vector database contract, run the Document Shape Test first. Short flat corpus, sign it. Long structured corpus, you were about to buy the wrong tool for the job you actually have.
THE UNCOMFORTABLE QUESTION
Your vector database was never the standard. It was the first thing that worked, and the industry turned one working solution into an assumption nobody re-examined. How much of the last two years of your retrieval architecture was a decision, and how much was a default you inherited and never tested?
The unit of retrieval is the document's structure, not its similarity to a query.
Get More Articles Like This
Getting your AI agent setup right is just the start. I'm documenting every mistake, fix, and lesson learned as I build PhantomByte.
Subscribe to receive updates when we publish new content. No spam, just real lessons from the trenches.
Build Real AI Infrastructure
PhantomByte teaches you to build real AI infrastructure yourself: local AI stacks, autonomous agents, multi-agent orchestration, web scraping, and custom tools. Step-by-step PDF tutorials you download, follow, and deploy. No subscriptions. No fluff. Just skills that ship.
