How Mafin 2.5 Uses PageIndex for Hierarchical Financial RAG

Mafin 2.5 leverages PageIndex as its core vector-less retrieval layer, using hierarchical tree structures instead of embeddings to achieve 98.7% accuracy on financial question-answering benchmarks.

The relationship between Mafin 2.5 and PageIndex represents a fundamental shift from traditional vector-based retrieval to hierarchical, reasoning-driven RAG architectures. While PageIndex provides the foundational document parsing and tree construction, Mafin 2.5 implements the domain-specific financial reasoning layer that navigates this structure to synthesize precise, cited answers.

What Is PageIndex?

PageIndex is an open-source library that transforms documents into vector-less hierarchical tree structures. According to the source code in pageindex/page_index.py, the PageIndex class parses PDFs or Markdown files into a navigable tree where each node contains:

  • LLM-generated summaries of the section content
  • Unique node IDs for precise referencing
  • Page ranges for accurate citations

This architecture enables language models to navigate documents similarly to human experts, moving through sections and subsections rather than relying on flat semantic similarity searches.

Understanding Mafin 2.5

Mafin 2.5 (officially released as the successor to Mafin 2) is a reasoning-based retrieval-augmented generation system specifically engineered for financial question answering. Unlike generic RAG implementations that depend solely on vector embeddings, Mafin 2.5 employs a reasoning engine that queries hierarchical document structures to select contextually relevant information.

As documented in the repository's README.md at line 73, this integration achieves state-of-the-art 98.7% accuracy on the FinanceBench benchmark, demonstrating that tree-based retrieval can outperform traditional embedding approaches for complex domain-specific queries.

Technical Integration: How Mafin 2.5 Leverages PageIndex

The relationship between Mafin 2.5 and PageIndex operates across three distinct technical layers: tree construction, hierarchical retrieval, and answer synthesis.

Building the Hierarchical Index

Mafin 2.5 utilizes PageIndex to parse source documents into navigable trees. The entry point run_pageindex.py demonstrates how the system invokes the PageIndex class:


# run_pageindex.py – generate the hierarchical index

from pageindex.page_index import PageIndex

pdf_path = "tests/pdfs/2023-annual-report.pdf"
index = PageIndex(pdf_path=pdf_path)
tree = index.build_tree()
print(tree)          # JSON-like tree ready for retrieval

This tree structure stores each document section with metadata including LLM-generated summaries and page references, creating the foundation for reasoning-based retrieval.

Reasoning Over the Tree Structure

Rather than performing flat vector similarity searches, Mafin 2.5 queries the PageIndex hierarchical structure. The pageindex/utils.py file contains helper functions such as retrieve_nodes that enable tree navigation:

from pageindex.utils import retrieve_nodes

query = "What was the company's net profit in 2023?"
relevant = retrieve_nodes(tree, query, top_k=3)
for node in relevant:
    print(node["title"], node["summary"])

Mafin 2.5 extends this basic retrieval with domain-specific financial reasoning, selecting nodes based on contextual relevance to complex multi-hop questions rather than simple keyword matching.

Answer Synthesis with Structured Context

The final integration layer involves synthesizing answers from retrieved nodes. Mafin 2.5 implements specialized prompting strategies that leverage PageIndex's structured summaries:

import openai

def answer_with_mafin(query, nodes):
    context = "\n".join([n["summary"] for n in nodes])
    prompt = f"""You are a financial analyst. Answer the question using ONLY the information below.
    
    Question: {query}
    
    Context:
    {context}
    
    Provide a concise answer and cite page numbers."""
    response = openai.ChatCompletion.create(
        model="gpt-4o-2024-11-20",
        messages=[{"role": "user", "content": prompt}]
    )
    return response.choices[0].message.content

print(answer_with_mafin(query, relevant))

This approach, documented in README.md at line 211, utilizes PageIndex's hierarchical structure to provide cited, accurate financial answers without relying on vector embeddings.

Key Implementation Files

The relationship between Mafin 2.5 and PageIndex relies on several critical source files:

  • run_pageindex.py: CLI entry point that builds hierarchical indexes from PDF or Markdown inputs.
  • pageindex/page_index.py: Core PageIndex class implementing document parsing, tree construction, and LLM-based summarization.
  • pageindex/utils.py: Utility functions for node retrieval, similarity scoring, and tree navigation used by Mafin 2.5's reasoning engine.
  • cookbook/pageindex_RAG_simple.ipynb: Jupyter notebook demonstrating vector-less RAG workflows similar to Mafin 2.5's architecture.

Performance Impact of the Integration

According to the repository's README.md, the combination of PageIndex's hierarchical retrieval and Mafin 2.5's reasoning capabilities achieves 98.7% accuracy on the FinanceBench benchmark. This performance demonstrates that tree-based retrieval with structured summaries can outperform traditional embedding-based approaches for complex domain-specific questions requiring precise citations and multi-hop reasoning.

Summary

  • PageIndex provides the foundational vector-less hierarchical indexing system that parses documents into navigable tree structures with LLM-generated summaries and page references.
  • Mafin 2.5 implements the domain-specific financial reasoning layer that queries this tree structure to retrieve relevant context and synthesize accurate, cited answers.
  • The integration achieves state-of-the-art performance (98.7% on FinanceBench) by replacing flat vector similarity with hierarchical reasoning-based retrieval.
  • Key implementation files include pageindex/page_index.py for tree construction and pageindex/utils.py for retrieval logic.

Frequently Asked Questions

Is PageIndex exclusively designed for financial documents?

No. While Mafin 2.5 specializes in financial question answering, PageIndex is a general-purpose document indexing library. The PageIndex class in pageindex/page_index.py can process any PDF or Markdown file, making it suitable for legal, medical, or technical documentation that benefits from hierarchical navigation and structured summarization.

Does Mafin 2.5 rely on vector embeddings for retrieval?

According to the source code analysis, Mafin 2.5 relies on PageIndex's vector-less hierarchical representation rather than traditional embeddings. Instead of embedding document chunks and performing similarity searches, the system uses the tree structure with LLM-generated summaries to navigate documents. This approach eliminates the need for vector databases while improving citation accuracy and reasoning capabilities.

Can I build a RAG system using PageIndex without Mafin 2.5?

Yes. The repository includes cookbook/pageindex_RAG_simple.ipynb, which demonstrates how to build a standalone vector-less RAG pipeline using PageIndex. You can use the PageIndex class to build document trees and the retrieval utilities in pageindex/utils.py to fetch relevant nodes for any domain-specific application without requiring Mafin 2.5's financial reasoning layer.

What specific advantages does the Mafin 2.5 and PageIndex integration provide over traditional RAG?

The integration achieves 98.7% accuracy on FinanceBench by combining PageIndex's hierarchical document representation with Mafin 2.5's reasoning capabilities. Traditional RAG systems rely on flat vector similarity, which can miss contextual relationships between sections or retrieve irrelevant chunks. The tree-based approach preserves document structure, enables precise page citations, and allows the reasoning engine to perform multi-hop navigation through related sections, resulting in more accurate and verifiable answers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →