How to Query the Knowledge Graph Using Natural Language in code-graph-rag

The code-graph-rag CLI translates natural-language questions into validated Cypher queries via an LLM-powered agent, executes them against Memgraph, and returns formatted results.

This guide explains how to query your codebase knowledge graph using plain English in the vitali87/code-graph-rag project. Whether you use the command line or the Python API, the system automatically generates safe, read-only Cypher queries and displays structured results.


CLI Method: Single Natural-Language Query

The fastest way to query is through the --ask-agent flag of the start command.

Command Structure

cgr start --repo-path /path/to/project \
          --ask-agent "Your natural language question here"

What Happens Under the Hood

  1. cli.py parses --ask-agent and calls main_single_query
  2. The RAG orchestrator (built in services/llm.py) loads the query tool
  3. query_codebase_knowledge_graph translates your question to Cypher using CypherGenerator
  4. The generated Cypher is validated and executed via MemgraphIngestor.fetch_all
  5. Results render as a Rich table in your terminal

Example Output

cgr start --repo-path ./my-project \
          --ask-agent "Find all functions that call each other"

╭─────────────────────────────────────╮
│ Query Results                        │
├───────────────┬─────────────────────┤
│ function_name │ caller_count        │
├───────────────┼─────────────────────┤
│ foo           │ 3                   │
│ bar           │ 1                   │
╰───────────────┴─────────────────────╯


Python API Method: Programmatic Natural-Language Queries

For integration into scripts or custom workflows, import the components directly.

Building the Query Pipeline

from codebase_rag.services.llm import CypherGenerator, create_rag_orchestrator
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings

# 1️⃣  Connect to Memgraph

ingestor = MemgraphIngestor(
    host=settings.MEMGRAPH_HOST,
    port=settings.MEMGRAPH_PORT,
    batch_size=500,
)

# 2️⃣  Create the Cypher generator

cypher_gen = CypherGenerator(active_projects=None)

# 3️⃣  Build the query tool

query_tool = create_query_tool(ingestor, cypher_gen)

# 4️⃣  Assemble the orchestrator

agent, _ = create_rag_orchestrator(tools=[query_tool])

# 5️⃣  Execute natural-language query

result = await agent.run("Show me all classes that implement the Repository interface")
print(result.output)

The agent.run() call triggers the full pipeline: NL → Cypher generation → execution → formatted output.


Direct Tool Invocation: Bypassing the Agent

For lower-level control, call the tool function directly without the agent loop.

import asyncio
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.services.graph_service import MemgraphIngestor

async def run_query():
    ingestor = MemgraphIngestor(host="localhost", port=7687)
    cypher_gen = CypherGenerator()
    tool = create_query_tool(ingestor, cypher_gen)

    # Execute directly — no agent overhead

    data = await tool.function("List all public functions in module user.auth")
    
    print(data.summary)      # Human-readable summary

    for row in data.results: # Raw Memgraph rows

        print(row)

asyncio.run(run_query())

This returns a QueryGraphData object containing:

  • cypher — the generated query string
  • results — list of row dictionaries
  • summary — formatted description of findings

How Natural Language Becomes a Cypher Query

The translation happens in services/llm.py. Understanding this flow helps you write more effective questions.

Step Function Location Purpose
1 CypherGenerator.generate services/llm.py:34 Sends NL question to LLM, receives raw Cypher
2 _clean_cypher_response services/llm.py Strips markdown fences and commentary
3 _validate_cypher_read_only services/llm.py Blocks CREATE, DELETE, SET, MERGE
4 _validate_no_unbounded_paths services/llm.py Prevents expensive unbounded traversals
5 _validate_call_procedures services/llm.py Restricts dangerous procedures

Only after all validations pass does MemgraphIngestor.fetch_all (in services/graph_service.py:47) execute the query with a configured memory limit.


Key Source Files for Natural-Language Queries

Purpose File Line
CLI entry with --ask-agent codebase_rag/cli.py 8
Orchestrator & CypherGenerator codebase_rag/services/llm.py 34, 56
Query tool definition codebase_rag/tools/codebase_query.py 30
Memgraph execution codebase_rag/services/graph_service.py 47
Interactive agent loop codebase_rag/main.py ~1650 (_run_agent_response_loop)

Writing Effective Natural-Language Queries

The LLM generates better Cypher when you include:

  • Specific entity types: "functions", "classes", "modules", "interfaces"
  • Relationship directions: "called by", "implements", "imports from"
  • Scope constraints: "in module X", "in file Y.py", "public methods only"
Less effective More effective
"Show me auth stuff" "List all functions in module auth that validate tokens"
"Find errors" "Find all classes that inherit from BaseException in the handlers package"
"What's connected to user?" "Show all functions that call methods on the User class, two levels deep"

Summary

  • CLI single query: Use cgr start --ask-agent "..." for quick terminal-based natural-language queries
  • Python agent: Import create_rag_orchestrator with create_query_tool for async programmatic access
  • Direct tool call: Use tool.function() directly when you need raw QueryGraphData without agent overhead
  • Safety guarantees: All Cypher is validated as read-only, bounded, and procedure-safe before execution against Memgraph

Frequently Asked Questions

What LLM does code-graph-rag use for natural-language to Cypher translation?

The CypherGenerator class in services/llm.py uses whichever model is configured in your environment (OpenAI, Anthropic, or local via LiteLLM). The specific provider and model name are read from settings, not hardcoded, so you can switch models without code changes.

Can I modify or extend the safety validations on generated Cypher?

Yes. The validation methods _validate_cypher_read_only, _validate_no_unbounded_paths, and _validate_call_procedures in services/llm.py are implemented as private methods on CypherGenerator. You can subclass CypherGenerator and override these methods, or add additional validators in generate() before the query reaches MemgraphIngestor.

What happens if the LLM generates invalid Cypher syntax?

The fetch_all method in services/graph_service.py catches execution errors and propagates them as exceptions. The tool wrapper in codebase_query.py logs the failed Cypher and returns an error message in the QueryGraphData summary field, allowing the agent to report the failure rather than crashing.

Is there a way to see the generated Cypher before it runs?

The QueryGraphData object returned by the tool includes the cypher attribute containing the exact query string sent to Memgraph. When using the CLI, this appears in verbose logs; when using the Python API, access it via data.cypher on the result object.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →