# How to Query the Knowledge Graph Using Natural Language in code-graph-rag

> Learn to query the knowledge graph using natural language with code-graph-rag. This tool translates your questions into Cypher queries for Memgraph, returning formatted results.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-20

---

**The code-graph-rag CLI translates natural-language questions into validated Cypher queries via an LLM-powered agent, executes them against Memgraph, and returns formatted results.**

This guide explains how to query your codebase knowledge graph using plain English in the `vitali87/code-graph-rag` project. Whether you use the command line or the Python API, the system automatically generates safe, read-only Cypher queries and displays structured results.

---

## CLI Method: Single Natural-Language Query

The fastest way to query is through the `--ask-agent` flag of the `start` command.

### Command Structure

```bash
cgr start --repo-path /path/to/project \
          --ask-agent "Your natural language question here"

```

### What Happens Under the Hood

1. **[`cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/cli.py)** parses `--ask-agent` and calls `main_single_query`
2. The **RAG orchestrator** (built in [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py)) loads the query tool
3. **`query_codebase_knowledge_graph`** translates your question to Cypher using `CypherGenerator`
4. The generated Cypher is validated and executed via `MemgraphIngestor.fetch_all`
5. Results render as a **Rich table** in your terminal

### Example Output

```bash
cgr start --repo-path ./my-project \
          --ask-agent "Find all functions that call each other"

```

```

╭─────────────────────────────────────╮
│ Query Results                        │
├───────────────┬─────────────────────┤
│ function_name │ caller_count        │
├───────────────┼─────────────────────┤
│ foo           │ 3                   │
│ bar           │ 1                   │
╰───────────────┴─────────────────────╯

```

---

## Python API Method: Programmatic Natural-Language Queries

For integration into scripts or custom workflows, import the components directly.

### Building the Query Pipeline

```python
from codebase_rag.services.llm import CypherGenerator, create_rag_orchestrator
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.graph_service import MemgraphIngestor
from codebase_rag.config import settings

# 1️⃣  Connect to Memgraph

ingestor = MemgraphIngestor(
    host=settings.MEMGRAPH_HOST,
    port=settings.MEMGRAPH_PORT,
    batch_size=500,
)

# 2️⃣  Create the Cypher generator

cypher_gen = CypherGenerator(active_projects=None)

# 3️⃣  Build the query tool

query_tool = create_query_tool(ingestor, cypher_gen)

# 4️⃣  Assemble the orchestrator

agent, _ = create_rag_orchestrator(tools=[query_tool])

# 5️⃣  Execute natural-language query

result = await agent.run("Show me all classes that implement the Repository interface")
print(result.output)

```

The `agent.run()` call triggers the full pipeline: NL → Cypher generation → execution → formatted output.

---

## Direct Tool Invocation: Bypassing the Agent

For lower-level control, call the tool function directly without the agent loop.

```python
import asyncio
from codebase_rag.tools.codebase_query import create_query_tool
from codebase_rag.services.llm import CypherGenerator
from codebase_rag.services.graph_service import MemgraphIngestor

async def run_query():
    ingestor = MemgraphIngestor(host="localhost", port=7687)
    cypher_gen = CypherGenerator()
    tool = create_query_tool(ingestor, cypher_gen)

    # Execute directly — no agent overhead

    data = await tool.function("List all public functions in module user.auth")
    
    print(data.summary)      # Human-readable summary

    for row in data.results: # Raw Memgraph rows

        print(row)

asyncio.run(run_query())

```

This returns a `QueryGraphData` object containing:
- **`cypher`** — the generated query string
- **`results`** — list of row dictionaries
- **`summary`** — formatted description of findings

---

## How Natural Language Becomes a Cypher Query

The translation happens in [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py). Understanding this flow helps you write more effective questions.

| Step | Function | Location | Purpose |
|------|----------|----------|---------|
| 1 | `CypherGenerator.generate` | `services/llm.py:34` | Sends NL question to LLM, receives raw Cypher |
| 2 | `_clean_cypher_response` | [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) | Strips markdown fences and commentary |
| 3 | `_validate_cypher_read_only` | [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) | Blocks `CREATE`, `DELETE`, `SET`, `MERGE` |
| 4 | `_validate_no_unbounded_paths` | [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) | Prevents expensive unbounded traversals |
| 5 | `_validate_call_procedures` | [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) | Restricts dangerous procedures |

Only after all validations pass does `MemgraphIngestor.fetch_all` (in `services/graph_service.py:47`) execute the query with a configured memory limit.

---

## Key Source Files for Natural-Language Queries

| Purpose | File | Line |
|---------|------|------|
| CLI entry with `--ask-agent` | [`codebase_rag/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cli.py) | 8 |
| Orchestrator & `CypherGenerator` | [`codebase_rag/services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/llm.py) | 34, 56 |
| Query tool definition | [`codebase_rag/tools/codebase_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/codebase_query.py) | 30 |
| Memgraph execution | [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) | 47 |
| Interactive agent loop | [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) | ~1650 (`_run_agent_response_loop`) |

---

## Writing Effective Natural-Language Queries

The LLM generates better Cypher when you include:

- **Specific entity types**: "functions", "classes", "modules", "interfaces"
- **Relationship directions**: "called by", "implements", "imports from"
- **Scope constraints**: "in module X", "in file Y.py", "public methods only"

| Less effective | More effective |
|--------------|---------------|
| "Show me auth stuff" | "List all functions in module auth that validate tokens" |
| "Find errors" | "Find all classes that inherit from BaseException in the handlers package" |
| "What's connected to user?" | "Show all functions that call methods on the User class, two levels deep" |

---

## Summary

- **CLI single query**: Use `cgr start --ask-agent "..."` for quick terminal-based natural-language queries
- **Python agent**: Import `create_rag_orchestrator` with `create_query_tool` for async programmatic access
- **Direct tool call**: Use `tool.function()` directly when you need raw `QueryGraphData` without agent overhead
- **Safety guarantees**: All Cypher is validated as read-only, bounded, and procedure-safe before execution against Memgraph

---

## Frequently Asked Questions

### What LLM does code-graph-rag use for natural-language to Cypher translation?

The `CypherGenerator` class in [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) uses whichever model is configured in your environment (OpenAI, Anthropic, or local via LiteLLM). The specific provider and model name are read from settings, not hardcoded, so you can switch models without code changes.

### Can I modify or extend the safety validations on generated Cypher?

Yes. The validation methods `_validate_cypher_read_only`, `_validate_no_unbounded_paths`, and `_validate_call_procedures` in [`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py) are implemented as private methods on `CypherGenerator`. You can subclass `CypherGenerator` and override these methods, or add additional validators in `generate()` before the query reaches `MemgraphIngestor`.

### What happens if the LLM generates invalid Cypher syntax?

The `fetch_all` method in [`services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/graph_service.py) catches execution errors and propagates them as exceptions. The tool wrapper in [`codebase_query.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_query.py) logs the failed Cypher and returns an error message in the `QueryGraphData` summary field, allowing the agent to report the failure rather than crashing.

### Is there a way to see the generated Cypher before it runs?

The `QueryGraphData` object returned by the tool includes the `cypher` attribute containing the exact query string sent to Memgraph. When using the CLI, this appears in verbose logs; when using the Python API, access it via `data.cypher` on the result object.