# How to Query Code Structure Using Natural Language With Code-Graph-RAG

> Query code structure with natural language using Code-Graph-RAG. Transform English questions into safe Cypher queries against your codebase graph database for quick insights.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Code-Graph-RAG converts plain-English questions about a codebase into safe, read-only Cypher queries that run against a graph database of parsed AST nodes, returning structured results without manual query writing.**

The `vitali87/code-graph-rag` open-source project bridges natural language and code analysis by combining Tree-sitter parsing, LLM-based query generation, and Memgraph storage. This article breaks down exactly how the four-layer pipeline works, where the critical code lives, and how to use it from both CLI and Python.

## The Four-Layer Architecture

Code-Graph-RAG processes every natural language query through four sequential layers. Understanding these layers helps you debug failures, customize behavior, and extend the system.

### Layer 1: Parsing & Graph Ingestion

Tree-sitter parses every source file—Python, TypeScript, Rust, and more—and extracts functions, classes, modules, and their relationships. These become nodes and edges in Memgraph.

- **Implementation**: [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py)
- **Output**: A labeled property graph with node types like `:Class`, `:Function`, `:Interface`, `:Module`

The graph schema is documented in [`docs/architecture/graph-schema.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/graph-schema.md) and constrains what queries the LLM can validly generate.

### Layer 2: Natural-Language → Cypher Translation

The `CypherGenerator` class in [`codebase_rag/services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/llm.py) (lines 136-138) wraps your configured LLM and formats a prompt that instructs the model to output **read-only Cypher only**. The method signature:

```python
cypher_query = await cypher_gen.generate(natural_language_query)

```

Prompt engineering ensures the generated query respects the schema: node labels, relationship types, and property names must match what was ingested in Layer 1.

### Layer 3: Query Execution & Safety Enforcement

[`codebase_rag/tools/codebase_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/codebase_query.py) contains `create_query_tool()`, which builds a callable `Tool` whose function is `query_codebase_knowledge_graph()`. Before execution, multiple safety layers run:

- **`requires_project_evidence()`** (lines 74-84): Validates that the Cypher returns at least one `n.qualified_name` field, enabling project-level scoping
- **Regex filters** (`_PROJECTED_QUALIFIED_NAME_RE`, `_AGGREGATED_ENTITY_RE`): Reject queries that could aggregate or leak data across projects

Only then does the query execute:

```python
results = await asyncio.to_thread(ingestor.fetch_all, cypher_query)

```

### Layer 4: Result Presentation

Results flow through `scope_rows_to_project()` to discard cross-project rows, then `truncate_results_by_tokens()` to respect the token budget from [`pyproject.toml`](https://github.com/vitali87/code-graph-rag/blob/main/pyproject.toml). The final `QueryGraphData` object contains:

- `query_used`: The original Cypher (for auditability)
- `results`: Filtered, truncated row data
- Human-readable summary

The CLI renders this with Rich tables and collapsible Cypher panels.

## Querying From the Interactive CLI

Start a session pointing at any repository:

```bash
cgr start --repo-path /path/to/repo

```

Then ask structural questions directly:

```text
> Find all classes that implement the Repository interface

```

Behind the scenes, the agent invokes the `query_graph` tool defined in [`codebase_rag/tools/tool_descriptions.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/tool_descriptions.py) (lines 25-31). The generated Cypher might look like:

```cypher
MATCH (c:Class)-[:IMPLEMENTS]->(i:Interface {name: 'Repository'})
RETURN c.qualified_name AS class_name, c.file AS file_path

```

The CLI displays results in a formatted table with the Cypher available for inspection.

## Programmatic Querying With the Python SDK

For integration into applications or notebooks, instantiate components directly:

```python
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.main import connect_memgraph

# Initialize once per session

ingestor = connect_memgraph()
cypher_gen = CypherGenerator()
tool = create_query_tool(
    ingestor, 
    cypher_gen, 
    project_name="myproj"  # Enables automatic scoping

)

# Execute natural language queries

answer = await tool.function(
    natural_language_query="Show me all async functions"
)

print(answer.query_used)   # Generated Cypher

print(answer.results)      # List[Dict] of matching entities

```

The `project_name` parameter activates filtering in `scope_rows_to_project()`, ensuring results belong only to the specified codebase.

## Deterministic MCP Queries Without LLM Latency

When you know exact qualified names, bypass the LLM entirely using deterministic tools in [`codebase_rag/graph_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_query.py):

```bash
cgr graph callers myproj.services.UserService.create_user --project myproj

```

This invokes `graph_query.callers()` (lines 40-50), which constructs static Cypher without model involvement—faster and fully reproducible. Other deterministic operations include `resolve` (lookup by qualified name) and dependency graph traversal.

## Safety Mechanisms for Production Use

Code-Graph-RAG treats LLM-generated Cypher as untrusted input. Key safeguards in [`codebase_rag/tools/codebase_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/codebase_query.py):

| Mechanism | Purpose | Location |
|-----------|---------|----------|
| Read-only enforcement | LLM prompt instructions + graph user permissions prevent writes | [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) prompt template |
| Project scoping | `scope_rows_to_project()` filters post-query | Lines ~90-100 |
| Token truncation | `truncate_results_by_tokens()` caps LLM context window usage | Called before result return |
| Qualified name validation | `requires_project_evidence()` ensures scopability | Lines 74-84 |

These layers allow deployment in multi-tenant environments where codebases must remain isolated.

## Summary

- **Code-Graph-RAG** transforms natural language into Cypher through `CypherGenerator.generate()` in [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py)
- **Safety is layered**: LLM prompt constraints, regex validation, project scoping, and token limits protect against misuse
- **CLI entry point**: `cgr start --repo-path <path>` enables interactive sessions with Rich-formatted output
- **Programmatic access**: `create_query_tool()` returns an awaitable function for embedding in async applications
- **Deterministic fallback**: [`codebase_rag/graph_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_query.py) provides LLM-free operations when qualified names are known

## Frequently Asked Questions

### What database does Code-Graph-RAG require?

Code-Graph-RAG requires **Memgraph**, a native in-memory graph database compatible with Cypher. The `connect_memgraph()` function in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) establishes this connection, and all generated queries assume Memgraph's specific Cypher dialect.

### Can I use Code-Graph-RAG with repositories that aren't Python?

Yes. The Tree-sitter ingestion layer in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) handles multiple languages including TypeScript, Rust, and others. The graph schema normalizes language-specific constructs (e.g., Rust traits vs. TypeScript interfaces) into consistent node labels like `:Interface` or `:Class`.

### How does project scoping prevent data leakage?

The `requires_project_evidence()` function validates that every Cypher returns a `qualified_name` property. After execution, `scope_rows_to_project()` filters results where `qualified_name` doesn't start with the configured `project_name`. This ensures that even if the LLM generates overly broad queries, only the intended codebase's entities return.

### Is there a cost to using the interactive CLI versus deterministic commands?

Yes. Interactive natural language queries invoke the LLM (cost: latency + token usage), while deterministic commands like `cgr graph callers` use static Cypher generation in [`codebase_rag/graph_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_query.py) with zero LLM overhead. For production scripts or high-frequency operations, prefer deterministic methods.