How the RAG System Converts Natural Language to Cypher Queries: A Technical Deep Dive
The RAG system converts natural language to Cypher queries through a multi-step pipeline: an orchestrator prompt instructs an LLM agent to act as a translator, the CypherGenerator.generate method forwards the question to the model, and cleaning and validation utilities ensure safe, executable output before execution against Memgraph.
The vitali87/code-graph-rag repository implements a Retrieval-Augmented Generation (RAG) pipeline that lets developers query a code knowledge graph using plain English. This article explains how natural language to Cypher query conversion works under the hood, tracing the exact code paths and components responsible for translation, validation, and execution.
The Conversion Pipeline: 6 Steps from Question to Query
The transformation follows a strict chain of responsibility. Each component handles a specific concern, ensuring the final Cypher is syntactically valid, semantically safe, and optimized for the underlying graph database.
Step 1: Orchestrator Prompt Defines the Translator Role
The pipeline starts with build_cypher_system_prompt in codebase_rag/prompts.py. This function constructs the system message that tells the LLM it must act as a Cypher translator and output only raw query syntax.
# From codebase_rag/prompts.py, lines 54-62
def build_cypher_system_prompt() -> str:
return """
You are a Cypher query translator. Convert natural language questions
into valid Cypher queries for a Neo4j-compatible graph database.
Rules:
- Output ONLY the Cypher query
- Do not explain your reasoning
- Always end with a semicolon
- Use parameterized queries where possible
"""
This prompt engineering is critical: it constrains the model's output format and eliminates conversational filler that would break automated parsing.
Step 2: LLM Service Initializes the Agent
The CypherGenerator class in codebase_rag/services/llm.py boots the translation engine. Its __init__ method (lines 14-31) loads the active LLM provider configuration, selects the appropriate system prompt (local or remote), and instantiates a pydantic_ai.Agent primed for model calls.
# Conceptual initialization flow
from codebase_rag.services.llm import CypherGenerator
generator = CypherGenerator(
active_projects=["myproject"] # Filters generated queries to relevant codebases
)
The agent encapsulates provider-specific details (OpenAI, local models, etc.) behind a unified interface.
Step 3: Natural Language Reaches the LLM
The generate method (lines 34-46) receives the raw user question and forwards it to the underlying agent. This is the actual LLM invocation point where natural language enters the model.
# From codebase_rag/services/llm.py, simplified
async def generate(self, question: str) -> str:
"""Send question to LLM, receive raw Cypher output."""
result = await self.agent.run(prompt=question)
raw_output = result.content
return self._clean_cypher_response(raw_output)
The method is async to support concurrent query generation without blocking the orchestrator.
Step 4: Response Cleaning Strips Markdown Artifacts
LLMs frequently wrap code in Markdown fences or add explanatory text. The _clean_cypher_response method (lines 31-57) normalizes these outputs into pure Cypher.
The cleaner performs three operations:
- Removes
cypher` andfences - Strips leading/trailing whitespace and backticks
- Ensures a trailing semicolon for statement termination
# Example transformation
raw_llm_output = """
```cypher
MATCH (f:Function {name: 'process_data'})
RETURN f;
"""
cleaned = _clean_cypher_response(raw_llm_output)
### Step 5: Safety Validation Prevents Harmful Queries
Before execution, three validators enforce sandbox boundaries in [`codebase_rag/services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/llm.py) (lines 71-100):
| Validator | Purpose | Implementation |
|-----------|---------|---------------|
| `_validate_cypher_read_only` | Blocks `CREATE`, `DELETE`, `SET`, `MERGE`, `DROP` | Pattern matching against prohibited keywords |
| `_validate_no_unbounded_paths` | Prevents expensive unbounded traversals (`-[]->` without node labels or limits) | AST inspection for anonymous variable-length paths |
| `_validate_call_procedures` | Whitelists safe procedures (`db.*`, `gds.*` read-only) | Checks `CALL` statements against `ALLOWED_PROCEDURES` |
These validators reference constants defined in `codebase_rag/constants/*.py`, which catalog safe Cypher keywords and prohibited patterns.
### Step 6: Tool Invocation Bridges Generation and Execution
The `create_query_tool` function in [`codebase_rag/tools/codebase_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/codebase_query.py) (lines 30-48) wraps the entire pipeline into a callable `Tool` that the orchestrator can invoke. This tool:
1. Accepts natural language input
2. Routes to `CypherGenerator.generate`
3. Executes validated Cypher via `ingestor.fetch_all`
4. Formats and truncates results for human consumption
```python
# From codebase_rag/tools/codebase_query.py
async def query_codebase_knowledge_graph(question: str) -> QueryResult:
cypher = await cypher_generator.generate(question)
# Validation occurs inside generate()
try:
results = await ingestor.fetch_all(cypher)
# Truncation and formatting logic (lines 51-78)
return QueryResult(
query_used=cypher,
results=results[:100], # Hard limit prevents output flooding
summary=f"Found {len(results)} results"
)
except Exception as e:
return QueryResult(error=str(e))
Complete Usage Examples
Using the High-Level Tool Interface
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.services import QueryProtocol
from rich.console import Console
# Initialize components
cypher_gen = CypherGenerator(active_projects=["myproject"])
ingestor = QueryProtocol() # Concrete Memgraph connector
tool = create_query_tool(ingestor, cypher_gen, console=Console())
# Execute natural language query
result = await tool.function("Which functions call `process_data`?")
print(result.query_used) # Generated Cypher string
print(result.summary) # Human-readable result count
print(result.results) # Raw graph data
Direct Generator Access for Debugging
from codebase_rag.services.llm import CypherGenerator
gen = CypherGenerator()
cypher = await gen.generate("List all classes that inherit from `BaseHandler`")
print(cypher)
# Expected output:
# MATCH (c:Class)-[:INHERITS_FROM]->(b:Class {name: 'BaseHandler'})
# RETURN c.name AS name, c.path AS path LIMIT 50;
Key Architectural Decisions
Prompt engineering over fine-tuning: The system uses carefully crafted system prompts rather than custom-trained models, making it provider-agnostic and faster to iterate.
Validation at generation time: Safety checks run before database contact, preventing arbitrary code execution and expensive query patterns.
Tool abstraction: The create_query_tool wrapper allows the orchestrator to treat Cypher generation as just another capability in its tool suite, maintaining consistent interfaces across retrieval, analysis, and modification operations.
Summary
build_cypher_system_promptincodebase_rag/prompts.pyconstrains the LLM to output-only Cypher translationCypherGenerator.__init__andgenerateincodebase_rag/services/llm.pyhandle provider-agnostic LLM invocation_clean_cypher_responseremoves Markdown artifacts and normalizes syntax- Three validators (
_validate_cypher_read_only,_validate_no_unbounded_paths,_validate_call_procedures) enforce read-only, bounded, safe queries create_query_toolincodebase_rag/tools/codebase_query.pyexposes the pipeline as an orchestrator-callable interface- Constants modules centralize security policies and allowed operations
Frequently Asked Questions
What LLM providers does the CypherGenerator support?
The CypherGenerator uses pydantic_ai.Agent as its abstraction layer, which supports OpenAI, Anthropic, and local models via configurable providers. The provider selection happens at initialization based on environment configuration, allowing deployment flexibility without code changes.
How does the system prevent destructive Cypher operations?
Three layered validators block write operations before database execution. _validate_cypher_read_only pattern-matches against prohibited keywords (CREATE, DELETE, MERGE, SET, DROP). These checks run on the generated string prior to any ingestor.fetch_all call, ensuring the Memgraph instance remains read-only from the RAG interface.
What happens when the LLM returns malformed Cypher?
The _clean_cypher_response method handles common formatting errors including Markdown code fences, extra whitespace, and missing semicolons. If cleaning produces invalid syntax, the validation layer catches parsing errors and returns an error result through the tool interface rather than propagating exceptions to the orchestrator.
Can I customize the system prompt for domain-specific Cypher?
Yes. The build_cypher_system_prompt function and its local variant build_local_cypher_system_prompt in codebase_rag/prompts.py can be extended or replaced. Pass a custom prompt to CypherGenerator during initialization to specialize for specific graph schemas, naming conventions, or query optimization preferences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →