# How to Migrate from Existing Code Search Tools to Code-Graph-RAG: A Complete Guide

> Migrate from existing code search tools to Code-Graph-RAG. Replace text indexing with a single command to build a language-aware graph index. Query using Cypher or natural language.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: migration-guide
- Published: 2026-08-19

---

**To migrate from existing code search tools to Code-Graph-RAG, replace your static text indexing with a one-time `cgr start` command that builds a language-aware graph index, then query via Cypher or natural-language semantic search using the Python SDK or CLI.**

Traditional code search utilities like `grep`, `ripgrep`, and Sourcegraph rely on plain-text matching that lacks cross-file context and language semantics. **Code-Graph-RAG** (vitali87/code-graph-rag) introduces a graph-based index powered by Tree-sitter that captures functions, classes, and their call relationships—enabling semantic queries and automated code editing through a Memgraph backend. This guide provides the exact steps to migrate from legacy text-based search to this rich, queryable graph architecture.

## Why Migrate from Traditional Code Search to Code-Graph-RAG?

Traditional tools operate on regex and file paths, while Code-Graph-RAG constructs a traversable knowledge graph of your codebase. Here is the capability comparison:

| Feature | Traditional tools | Code-Graph-RAG |
|---------|-------------------|----------------|
| Exact-text search | ✅ | ✅ (via semantic layer) |
| Cross-file call graph | ❌ | ✅ |
| Multi-language AST parsing | ❌ | ✅ (Tree-sitter) |
| Natural-language query → Cypher | ❌ | ✅ (LLM-driven) |
| Semantic (embedding) search | ❌ | ✅ (UniXcoder) |
| Interactive code-editing | ❌ | ✅ (AST-guided patches) |

The architectural advantages are documented in the **Architecture Overview** at [`docs/architecture/overview.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/overview.md). The core systems include a **Multi-Language Parser** that uses Tree-sitter to build ASTs and extract symbols, and a **RAG System** (`codebase_rag/`) that translates natural language into executable Cypher queries.

## Step-by-Step Migration Workflow

### Install Code-Graph-RAG

Begin the migration by installing the package with semantic search extras:

```bash
pip install 'code-graph-rag[semantic]'

```

This command pulls in UniXcoder and vector-store dependencies required for embedding-based search. See the *Semantic Search* installation guide in [`docs/sdk/semantic-search.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/sdk/semantic-search.md) for backend-specific configuration.

### Replace Indexing with Graph Construction

Instead of running file-search scripts or building static indexes, invoke the parser once per repository:

```bash
cgr start --repo-path /path/to/your/project

```

Internally, this executes the Tree-sitter based parser located in `codebase_rag/parsers/` and persists the results to Memgraph via `cgr connect_memgraph`. This single command replaces your existing indexing pipeline with a rich, language-aware graph.

### Persist and Export Your Graph

For backup, CI pipelines, or offline analysis, export the graph to JSON:

```python
from cgr import export_graph_to_file, connect_memgraph

ingestor = connect_memgraph(batch_size=500)
export_graph_to_file(ingestor, "my_project_graph.json")

```

The exporter implementation resides in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) and uses the `export_graph_to_file` function. This allows you to version-control your code's structure independently of the live database.

### Replace Text Searches with Cypher and Semantic Queries

For exact matches previously handled by `grep`, use Cypher directly:

```cypher
MATCH (f:Function) WHERE f.name CONTAINS "parse" RETURN f

```

For conceptual queries, use the semantic layer leveraging UniXcoder embeddings:

```bash
cgr semantic-search "parse a JSON payload"

```

The semantic search system generates 768-dimensional vectors and queries the backend configured via `CGR_VECTOR_STORE_BACKEND` (Qdrant or Milvus), as detailed in the **Semantic Search** SDK documentation.

### Upgrade to the Interactive Chat Interface

Replace REPL-style grep sessions with the `cgr chat` CLI:

```bash
cgr chat

```

This interface supports commands like `/model` to switch models, `/help` for available operations, and automatically handles tool approvals for code edits. UI constants and prompt strings are defined in [`codebase_rag/constants/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/cli.py) (e.g., `UI_TOOL_APPROVAL` for approval prompts).

### Migrate Existing Scripts to the SDK

If you have Python scripts that currently wrap `grep` or parse file lists, refactor them to use the **Cypher Generator** ([`services/llm.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/llm.py)) and the **GraphLoader** class from [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py). This maintains programmatic access while gaining graph traversal capabilities.

## Architecture Overview for Migration

Understanding the component hierarchy ensures you map old workflows to the correct new subsystems:

| Layer | Responsibility | Key Implementation |
|-------|----------------|--------------------|
| **Parser** | Tree-sitter drives language-agnostic AST extraction. | `codebase_rag/parsers/` – uses `tree_sitter` bindings. |
| **Graph Ingestion** | Translates AST nodes/relationships into Memgraph entities. | [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) and [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py). |
| **Embedding / Semantic Index** | Generates 768-dim vectors with UniXcoder. | [`codebase_rag/services/embedding.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/embedding.py); configure via `CGR_VECTOR_STORE_BACKEND`. |
| **Query Engine** | LLM (via `pydantic_ai`) converts NL → Cypher. | [`services/cypher_generator.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/cypher_generator.py); orchestrated by [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py). |
| **CLI / UI** | Rich-based status bar, prompts, and diff rendering. | [`codebase_rag/constants/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/cli.py). |
| **Persistence** | Export / import JSON for offline analysis. | `export_graph_to_file` in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py). |

All components are wired together by the **AppContext** singleton ([`codebase_rag/models.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/models.py)), which propagates configuration (`settings`) and runtime state throughout the application.

## Code Examples: From Legacy Tools to Code-Graph-RAG

### Replacing grep with Semantic Search

Convert legacy text searches to embedding-based retrieval:

```python
from cgr import embed_code, semantic_search

# Generate embedding for a code snippet (usually done once during indexing)

snippet = """
def validate_user(token: str) -> bool:
    # verify JWT token

    ...
"""
embedding = embed_code(snippet)          # uses UniXcoder under the hood

print(f"Embedding size: {len(embedding)}")   # → 768

# Perform semantic search

results = semantic_search("jwt validation function")
for node_id, score in results:
    print(f"Node {node_id} – similarity {score:.2f}")

```

The `embed_code` function resides in [`codebase_rag/services/embedding.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/embedding.py), while `semantic_search` reaches the vector store configured via environment variables.

### Converting Text Searches to Cypher Queries

Migrate shell scripts that invoke `grep` to Python using the Graph SDK:

```python
from cgr import GraphLoader, connect_memgraph

# Load an existing exported graph

loader = GraphLoader("my_project_graph.json")
loader.load()

# Old grep:   grep -R "def .*parse.*" .

# New Cypher:

query = """
MATCH (f:Function)
WHERE f.name =~ '(?i).*parse.*'
RETURN f.name AS function, f.file AS file
ORDER BY f.name
"""
results = loader.graph_service.run_cypher(query)
for row in results:
    print(f"{row['function']}  –  {row['file']}")

```

The `GraphLoader` class in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) handles JSON deserialization, while `graph_service` wraps the Memgraph client for Cypher execution.

### Automating Code Edits via RAG

Use the interactive CLI for AI-guided refactoring:

```bash
cgr chat
> replace the logging call in src/utils.py with logger.info(...)

```

The system executes three steps defined in [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py):
1. **Generate a diff** via `_print_unified_diff` showing the proposed change.
2. **Prompt for approval** using `_process_tool_approvals` (referencing `UI_TOOL_APPROVAL` from [`codebase_rag/constants/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/cli.py)).
3. **Apply the edit** via `FileEditor` ([`tools/file_editor.py`](https://github.com/vitali87/code-graph-rag/blob/main/tools/file_editor.py)).

## Key Files and Source References

| File | Role |
|------|------|
| [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py) | Central CLI entry point and interactive loops |
| [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) | JSON graph loading and export utilities |
| [`codebase_rag/constants/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/cli.py) | UI constants, approval prompts, and status messages |
| `codebase_rag/parsers/` | Tree-sitter AST extraction and multi-language parsing |
| [`codebase_rag/services/embedding.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/embedding.py) | UniXcoder embedding generation |
| [`services/cypher_generator.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/cypher_generator.py) | LLM-powered natural language to Cypher translation |
| [`docs/architecture/overview.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/architecture/overview.md) | High-level architecture documentation |
| [`docs/sdk/semantic-search.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/sdk/semantic-search.md) | Vector store configuration and semantic search usage |

## Summary

- Replace static indexing with a one-time `cgr start` command to build a language-aware graph in Memgraph.
- Query the graph using **Cypher** for exact structural matches or **semantic search** (UniXcoder embeddings) for conceptual queries.
- Use the **Python SDK** (`GraphLoader`, `semantic_search`) to refactor existing grep-based scripts.
- Leverage `cgr chat` for interactive, AI-assisted code editing with approval workflows.
- Export graphs to JSON via `export_graph_to_file` for backup and CI pipeline integration.

## Frequently Asked Questions

### How do I migrate existing grep-based scripts to Code-Graph-RAG?

Wrap your existing logic using the `GraphLoader` class from [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) and the Cypher execution methods in [`services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/services/graph_service.py). Replace regex patterns with `MATCH` clauses that query function names, file paths, and relationships. This preserves your automation while adding cross-file context.

### Can I migrate incrementally without disrupting existing workflows?

Yes. Code-Graph-RAG can coexist with traditional tools during the transition. Run `cgr start` to build the initial graph, then gradually migrate specific queries to the new system while keeping legacy scripts operational. The JSON export feature allows you to maintain offline snapshots without database dependency.

### Which programming languages are supported when migrating my codebase?

The system uses Tree-sitter for parsing, supporting any language with a Tree-sitter grammar. The parser implementation in `codebase_rag/parsers/` extracts symbols and relationships uniformly across languages, making the migration path consistent regardless of whether you are searching Python, JavaScript, Go, or Rust codebases.

### How is cross-file call graph analysis handled during migration?

Traditional tools require manual navigation between files. After migration, the graph schema directly stores `CALLS` relationships between functions across files. Query these using Cypher: `MATCH (caller:Function)-[:CALLS]->(callee:Function) RETURN caller, callee`. This is powered by the graph ingestion layer in [`codebase_rag/graph_updater.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_updater.py) and the AST pattern matching defined in `codebase_rag/parsers/ast_grep_patterns/`.