How to Query the Knowledge Graph Using Natural Language in code-graph-rag
The code-graph-rag CLI translates natural-language questions into validated Cypher queries via an LLM-powered agent, executes them against Memgraph, and returns formatted results.
This guide explains how to query your codebase knowledge graph using plain English in the vitali87/code-graph-rag project. Whether you use the command line or the Python API, the system automatically generates safe, read-only Cypher queries and displays structured results.
CLI Method: Single Natural-Language Query
The fastest way to query is through the --ask-agent flag of the start command.
Command Structure
cgr start --repo-path /path/to/project \
--ask-agent "Your natural language question here"
What Happens Under the Hood
cli.pyparses--ask-agentand callsmain_single_query- The RAG orchestrator (built in
services/llm.py) loads the query tool query_codebase_knowledge_graphtranslates your question to Cypher usingCypherGenerator- The generated Cypher is validated and executed via
MemgraphIngestor.fetch_all - Results render as a Rich table in your terminal
Example Output
cgr start --repo-path ./my-project \
--ask-agent "Find all functions that call each other"
╭─────────────────────────────────────╮
│ Query Results │
├───────────────┬─────────────────────┤
│ function_name │ caller_count │
├───────────────┼─────────────────────┤
│ foo │ 3 │
│ bar │ 1 │
╰───────────────┴─────────────────────╯
Python API Method: Programmatic Natural-Language Queries
For integration into scripts or custom workflows, import the components directly.
Building the Query Pipeline
from codebase_rag.services.llm import CypherGenerator, create_rag_orchestrator
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings
# 1️⃣ Connect to Memgraph
ingestor = MemgraphIngestor(
host=settings.MEMGRAPH_HOST,
port=settings.MEMGRAPH_PORT,
batch_size=500,
)
# 2️⃣ Create the Cypher generator
cypher_gen = CypherGenerator(active_projects=None)
# 3️⃣ Build the query tool
query_tool = create_query_tool(ingestor, cypher_gen)
# 4️⃣ Assemble the orchestrator
agent, _ = create_rag_orchestrator(tools=[query_tool])
# 5️⃣ Execute natural-language query
result = await agent.run("Show me all classes that implement the Repository interface")
print(result.output)
The agent.run() call triggers the full pipeline: NL → Cypher generation → execution → formatted output.
Direct Tool Invocation: Bypassing the Agent
For lower-level control, call the tool function directly without the agent loop.
import asyncio
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.services.graph_service import MemgraphIngestor
async def run_query():
ingestor = MemgraphIngestor(host="localhost", port=7687)
cypher_gen = CypherGenerator()
tool = create_query_tool(ingestor, cypher_gen)
# Execute directly — no agent overhead
data = await tool.function("List all public functions in module user.auth")
print(data.summary) # Human-readable summary
for row in data.results: # Raw Memgraph rows
print(row)
asyncio.run(run_query())
This returns a QueryGraphData object containing:
cypher— the generated query stringresults— list of row dictionariessummary— formatted description of findings
How Natural Language Becomes a Cypher Query
The translation happens in services/llm.py. Understanding this flow helps you write more effective questions.
| Step | Function | Location | Purpose |
|---|---|---|---|
| 1 | CypherGenerator.generate |
services/llm.py:34 |
Sends NL question to LLM, receives raw Cypher |
| 2 | _clean_cypher_response |
services/llm.py |
Strips markdown fences and commentary |
| 3 | _validate_cypher_read_only |
services/llm.py |
Blocks CREATE, DELETE, SET, MERGE |
| 4 | _validate_no_unbounded_paths |
services/llm.py |
Prevents expensive unbounded traversals |
| 5 | _validate_call_procedures |
services/llm.py |
Restricts dangerous procedures |
Only after all validations pass does MemgraphIngestor.fetch_all (in services/graph_service.py:47) execute the query with a configured memory limit.
Key Source Files for Natural-Language Queries
| Purpose | File | Line |
|---|---|---|
CLI entry with --ask-agent |
codebase_rag/cli.py |
8 |
Orchestrator & CypherGenerator |
codebase_rag/services/llm.py |
34, 56 |
| Query tool definition | codebase_rag/tools/codebase_query.py |
30 |
| Memgraph execution | codebase_rag/services/graph_service.py |
47 |
| Interactive agent loop | codebase_rag/main.py |
~1650 (_run_agent_response_loop) |
Writing Effective Natural-Language Queries
The LLM generates better Cypher when you include:
- Specific entity types: "functions", "classes", "modules", "interfaces"
- Relationship directions: "called by", "implements", "imports from"
- Scope constraints: "in module X", "in file Y.py", "public methods only"
| Less effective | More effective |
|---|---|
| "Show me auth stuff" | "List all functions in module auth that validate tokens" |
| "Find errors" | "Find all classes that inherit from BaseException in the handlers package" |
| "What's connected to user?" | "Show all functions that call methods on the User class, two levels deep" |
Summary
- CLI single query: Use
cgr start --ask-agent "..."for quick terminal-based natural-language queries - Python agent: Import
create_rag_orchestratorwithcreate_query_toolfor async programmatic access - Direct tool call: Use
tool.function()directly when you need rawQueryGraphDatawithout agent overhead - Safety guarantees: All Cypher is validated as read-only, bounded, and procedure-safe before execution against Memgraph
Frequently Asked Questions
What LLM does code-graph-rag use for natural-language to Cypher translation?
The CypherGenerator class in services/llm.py uses whichever model is configured in your environment (OpenAI, Anthropic, or local via LiteLLM). The specific provider and model name are read from settings, not hardcoded, so you can switch models without code changes.
Can I modify or extend the safety validations on generated Cypher?
Yes. The validation methods _validate_cypher_read_only, _validate_no_unbounded_paths, and _validate_call_procedures in services/llm.py are implemented as private methods on CypherGenerator. You can subclass CypherGenerator and override these methods, or add additional validators in generate() before the query reaches MemgraphIngestor.
What happens if the LLM generates invalid Cypher syntax?
The fetch_all method in services/graph_service.py catches execution errors and propagates them as exceptions. The tool wrapper in codebase_query.py logs the failed Cypher and returns an error message in the QueryGraphData summary field, allowing the agent to report the failure rather than crashing.
Is there a way to see the generated Cypher before it runs?
The QueryGraphData object returned by the tool includes the cypher attribute containing the exact query string sent to Memgraph. When using the CLI, this appears in verbose logs; when using the Python API, access it via data.cypher on the result object.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →