Graphify Benchmark: Graph Queries vs Full Corpus Reads Performance Comparison

Graphify's benchmark demonstrates that graph-based queries achieve 82% key-fact coverage while consuming only ~140K tokens per query, compared to 70.8% coverage and ~2.8M tokens for full corpus reads—delivering superior accuracy with 20x lower token costs.

Graphify-Labs/graphify evaluates retrieval performance across the ERPNext repository, a real-world Python codebase containing approximately one million lines of code. This analysis compares graph queries—leveraging deterministic AST extraction and graph expansion—against a naive full corpus read baseline using simple grep and file loading pipelines.

Code Intelligence Benchmark Methodology

The benchmark suite measures two complementary use-cases: conversational memory (LOCOMO/LongMemEval) and code intelligence on ERPNext. For the code-intelligence experiments documented in BENCHMARKS.md, the system compares Graphify's graph-expanded retrieval against a "search raw files" baseline implementing a simple grep plus read pipeline.

The ERPNext dataset spans 15 years of development history, growing from 3,000 nodes in 2011 to over 22,000 nodes in 2026. This temporal scale provides a realistic testbed for measuring how retrieval methods perform as codebases expand, as shown in Table "Temporal (15 years of ERPNext)" (lines 555-562).

Performance Results: Graph Queries vs Full Corpus Reads

The benchmark results reveal dramatic efficiency gains for graph-based retrieval:

Metric Graphify (Graph-Expand) Raw-File Grep + Read Baseline
Key-fact coverage 82% 70.8%
Tokens per query ~140K ~2.8M

Graphify's graph queries deliver approximately 11% higher coverage while using roughly 1/20th of the token budget. This efficiency stems from avoiding the need to load entire repository contents into the LLM context window.

Why Graph Queries Outperform Full Repository Reads

Three architectural advantages explain the performance gap between graph-based retrieval and full corpus reads.

Deterministic Graph Construction via AST Extraction

Graphify builds structured knowledge graphs using Tree-sitter to extract abstract syntax trees (ASTs), creating nodes for symbols, imports, and call relationships. This deterministic construction requires zero LLM credits during the indexing phase, unlike embedding-based approaches that consume tokens for every file. The graph structure captures semantic relationships that lexical search misses.

Hybrid Retrieval with Semantic Reranking

The retrieval pipeline in graphify/watch.py first expands queries over the graph to identify relevant nodes, then applies lightweight semantic reranking. This targeted approach retrieves only contextually necessary code snippets rather than entire files. Full corpus reads, by contrast, dump massive text volumes into the context window, degrading model performance and increasing costs.

Temporal Scalability Across 15 Years of Development

As shown in the benchmark's temporal analysis, simple lexical search hits a decreasing fraction of relevant facts as the codebase grows from 3K to 22K+ nodes. Graph-based retrieval maintains consistent coverage because relationship traversals scale with semantic complexity rather than file count.

Key Implementation Files

The benchmark implementation relies on several critical components:

  • BENCHMARKS.md – Contains the full experimental design, dataset descriptions, and quantitative results comparing graph queries versus full corpus reads.
  • crosstool/run.py – Executes the code-intelligence benchmark suite, invoking both Graphify and the raw-file baseline for direct comparison.
  • graphify/watch.py – Implements the core graph-expansion and retrieval logic that enables efficient query processing.
  • graphify/tree_html.py – Provides AST visualization utilities that demonstrate how graph structures enable precise targeting.

Practical Usage Example

The following code demonstrates how to query ERPNext using Graphify's efficient graph-based retrieval:

from graphify import GraphifyClient

client = GraphifyClient()
answer = client.query(
    repo="erpnext",
    question="How does the `sale_order` model calculate total amount?",
    max_tokens=140_000,
)

print(answer)

Running the same question against the raw-file baseline would require loading the entire repository into the model context, consuming approximately 2-3 million tokens while delivering lower factual accuracy.

Summary

  • Graphify's benchmark shows graph queries achieve 82% key-fact coverage compared to 70.8% for full corpus reads.
  • Graph-based retrieval uses approximately 140K tokens per query versus 2.8M tokens for grep-based baselines—a 20x reduction.
  • Deterministic AST extraction via Tree-sitter creates structured graphs without consuming LLM credits during indexing.
  • The hybrid retrieval architecture in graphify/watch.py enables scalable performance across 15 years of ERPNext development history.

Frequently Asked Questions

What dataset does Graphify use to compare graph queries against full corpus reads?

Graphify evaluates performance on ERPNext, a real-world Python repository containing approximately one million lines of code spanning 15 years of development (3,000 to 22,000+ nodes). This provides realistic conditions for testing retrieval accuracy and token efficiency at scale.

How does Graphify achieve 20x lower token usage than full repository reads?

Graphify's graphify/watch.py implementation uses deterministic graph expansion to retrieve only relevant symbol relationships and code snippets. Full corpus reads load entire files into the context window, consuming ~2.8M tokens compared to ~140K for targeted graph queries.

Why does graph-based retrieval maintain accuracy as codebases grow larger?

As demonstrated in the temporal analysis (lines 555-562 of BENCHMARKS.md), lexical search degrades as repositories scale because keyword matching becomes less precise. Graph queries traverse semantic relationships (imports, calls, inheritance) that remain stable regardless of total file count or repository age.

Where can I find the benchmark implementation comparing these approaches?

The benchmark harness resides in crosstool/run.py, which orchestrates side-by-side comparisons between Graphify and raw-file baselines. Detailed methodology and results tables are documented in BENCHMARKS.md at the repository root.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →