How to Query Code Structure Using Natural Language With Code-Graph-RAG

Code-Graph-RAG converts plain-English questions about a codebase into safe, read-only Cypher queries that run against a graph database of parsed AST nodes, returning structured results without manual query writing.

The vitali87/code-graph-rag open-source project bridges natural language and code analysis by combining Tree-sitter parsing, LLM-based query generation, and Memgraph storage. This article breaks down exactly how the four-layer pipeline works, where the critical code lives, and how to use it from both CLI and Python.

The Four-Layer Architecture

Code-Graph-RAG processes every natural language query through four sequential layers. Understanding these layers helps you debug failures, customize behavior, and extend the system.

Layer 1: Parsing & Graph Ingestion

Tree-sitter parses every source file—Python, TypeScript, Rust, and more—and extracts functions, classes, modules, and their relationships. These become nodes and edges in Memgraph.

The graph schema is documented in docs/architecture/graph-schema.md and constrains what queries the LLM can validly generate.

Layer 2: Natural-Language → Cypher Translation

The CypherGenerator class in codebase_rag/services/llm.py (lines 136-138) wraps your configured LLM and formats a prompt that instructs the model to output read-only Cypher only. The method signature:

cypher_query = await cypher_gen.generate(natural_language_query)

Prompt engineering ensures the generated query respects the schema: node labels, relationship types, and property names must match what was ingested in Layer 1.

Layer 3: Query Execution & Safety Enforcement

codebase_rag/tools/codebase_query.py contains create_query_tool(), which builds a callable Tool whose function is query_codebase_knowledge_graph(). Before execution, multiple safety layers run:

  • requires_project_evidence() (lines 74-84): Validates that the Cypher returns at least one n.qualified_name field, enabling project-level scoping
  • Regex filters (_PROJECTED_QUALIFIED_NAME_RE, _AGGREGATED_ENTITY_RE): Reject queries that could aggregate or leak data across projects

Only then does the query execute:

results = await asyncio.to_thread(ingestor.fetch_all, cypher_query)

Layer 4: Result Presentation

Results flow through scope_rows_to_project() to discard cross-project rows, then truncate_results_by_tokens() to respect the token budget from pyproject.toml. The final QueryGraphData object contains:

  • query_used: The original Cypher (for auditability)
  • results: Filtered, truncated row data
  • Human-readable summary

The CLI renders this with Rich tables and collapsible Cypher panels.

Querying From the Interactive CLI

Start a session pointing at any repository:

cgr start --repo-path /path/to/repo

Then ask structural questions directly:

> Find all classes that implement the Repository interface

Behind the scenes, the agent invokes the query_graph tool defined in codebase_rag/tools/tool_descriptions.py (lines 25-31). The generated Cypher might look like:

MATCH (c:Class)-[:IMPLEMENTS]->(i:Interface {name: 'Repository'})
RETURN c.qualified_name AS class_name, c.file AS file_path

The CLI displays results in a formatted table with the Cypher available for inspection.

Programmatic Querying With the Python SDK

For integration into applications or notebooks, instantiate components directly:

from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.main import connect_memgraph

# Initialize once per session

ingestor = connect_memgraph()
cypher_gen = CypherGenerator()
tool = create_query_tool(
    ingestor, 
    cypher_gen, 
    project_name="myproj"  # Enables automatic scoping

)

# Execute natural language queries

answer = await tool.function(
    natural_language_query="Show me all async functions"
)

print(answer.query_used)   # Generated Cypher

print(answer.results)      # List[Dict] of matching entities

The project_name parameter activates filtering in scope_rows_to_project(), ensuring results belong only to the specified codebase.

Deterministic MCP Queries Without LLM Latency

When you know exact qualified names, bypass the LLM entirely using deterministic tools in codebase_rag/graph_query.py:

cgr graph callers myproj.services.UserService.create_user --project myproj

This invokes graph_query.callers() (lines 40-50), which constructs static Cypher without model involvement—faster and fully reproducible. Other deterministic operations include resolve (lookup by qualified name) and dependency graph traversal.

Safety Mechanisms for Production Use

Code-Graph-RAG treats LLM-generated Cypher as untrusted input. Key safeguards in codebase_rag/tools/codebase_query.py:

Mechanism Purpose Location
Read-only enforcement LLM prompt instructions + graph user permissions prevent writes services/llm.py prompt template
Project scoping scope_rows_to_project() filters post-query Lines ~90-100
Token truncation truncate_results_by_tokens() caps LLM context window usage Called before result return
Qualified name validation requires_project_evidence() ensures scopability Lines 74-84

These layers allow deployment in multi-tenant environments where codebases must remain isolated.

Summary

  • Code-Graph-RAG transforms natural language into Cypher through CypherGenerator.generate() in services/llm.py
  • Safety is layered: LLM prompt constraints, regex validation, project scoping, and token limits protect against misuse
  • CLI entry point: cgr start --repo-path <path> enables interactive sessions with Rich-formatted output
  • Programmatic access: create_query_tool() returns an awaitable function for embedding in async applications
  • Deterministic fallback: codebase_rag/graph_query.py provides LLM-free operations when qualified names are known

Frequently Asked Questions

What database does Code-Graph-RAG require?

Code-Graph-RAG requires Memgraph, a native in-memory graph database compatible with Cypher. The connect_memgraph() function in codebase_rag/main.py establishes this connection, and all generated queries assume Memgraph's specific Cypher dialect.

Can I use Code-Graph-RAG with repositories that aren't Python?

Yes. The Tree-sitter ingestion layer in codebase_rag/graph_loader.py handles multiple languages including TypeScript, Rust, and others. The graph schema normalizes language-specific constructs (e.g., Rust traits vs. TypeScript interfaces) into consistent node labels like :Interface or :Class.

How does project scoping prevent data leakage?

The requires_project_evidence() function validates that every Cypher returns a qualified_name property. After execution, scope_rows_to_project() filters results where qualified_name doesn't start with the configured project_name. This ensures that even if the LLM generates overly broad queries, only the intended codebase's entities return.

Is there a cost to using the interactive CLI versus deterministic commands?

Yes. Interactive natural language queries invoke the LLM (cost: latency + token usage), while deterministic commands like cgr graph callers use static Cypher generation in codebase_rag/graph_query.py with zero LLM overhead. For production scripts or high-frequency operations, prefer deterministic methods.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →