How to Migrate from Existing Code Search Tools to Code-Graph-RAG: A Complete Guide

To migrate from existing code search tools to Code-Graph-RAG, replace your static text indexing with a one-time cgr start command that builds a language-aware graph index, then query via Cypher or natural-language semantic search using the Python SDK or CLI.

Traditional code search utilities like grep, ripgrep, and Sourcegraph rely on plain-text matching that lacks cross-file context and language semantics. Code-Graph-RAG (vitali87/code-graph-rag) introduces a graph-based index powered by Tree-sitter that captures functions, classes, and their call relationships—enabling semantic queries and automated code editing through a Memgraph backend. This guide provides the exact steps to migrate from legacy text-based search to this rich, queryable graph architecture.

Why Migrate from Traditional Code Search to Code-Graph-RAG?

Traditional tools operate on regex and file paths, while Code-Graph-RAG constructs a traversable knowledge graph of your codebase. Here is the capability comparison:

Feature Traditional tools Code-Graph-RAG
Exact-text search ✅ ✅ (via semantic layer)
Cross-file call graph ❌ ✅
Multi-language AST parsing ❌ ✅ (Tree-sitter)
Natural-language query → Cypher ❌ ✅ (LLM-driven)
Semantic (embedding) search ❌ ✅ (UniXcoder)
Interactive code-editing ❌ ✅ (AST-guided patches)

The architectural advantages are documented in the Architecture Overview at docs/architecture/overview.md. The core systems include a Multi-Language Parser that uses Tree-sitter to build ASTs and extract symbols, and a RAG System (codebase_rag/) that translates natural language into executable Cypher queries.

Step-by-Step Migration Workflow

Install Code-Graph-RAG

Begin the migration by installing the package with semantic search extras:

pip install 'code-graph-rag[semantic]'

This command pulls in UniXcoder and vector-store dependencies required for embedding-based search. See the Semantic Search installation guide in docs/sdk/semantic-search.md for backend-specific configuration.

Replace Indexing with Graph Construction

Instead of running file-search scripts or building static indexes, invoke the parser once per repository:

cgr start --repo-path /path/to/your/project

Internally, this executes the Tree-sitter based parser located in codebase_rag/parsers/ and persists the results to Memgraph via cgr connect_memgraph. This single command replaces your existing indexing pipeline with a rich, language-aware graph.

Persist and Export Your Graph

For backup, CI pipelines, or offline analysis, export the graph to JSON:

from cgr import export_graph_to_file, connect_memgraph

ingestor = connect_memgraph(batch_size=500)
export_graph_to_file(ingestor, "my_project_graph.json")

The exporter implementation resides in codebase_rag/main.py and uses the export_graph_to_file function. This allows you to version-control your code's structure independently of the live database.

Replace Text Searches with Cypher and Semantic Queries

For exact matches previously handled by grep, use Cypher directly:

MATCH (f:Function) WHERE f.name CONTAINS "parse" RETURN f

For conceptual queries, use the semantic layer leveraging UniXcoder embeddings:

cgr semantic-search "parse a JSON payload"

The semantic search system generates 768-dimensional vectors and queries the backend configured via CGR_VECTOR_STORE_BACKEND (Qdrant or Milvus), as detailed in the Semantic Search SDK documentation.

Upgrade to the Interactive Chat Interface

Replace REPL-style grep sessions with the cgr chat CLI:

cgr chat

This interface supports commands like /model to switch models, /help for available operations, and automatically handles tool approvals for code edits. UI constants and prompt strings are defined in codebase_rag/constants/cli.py (e.g., UI_TOOL_APPROVAL for approval prompts).

Migrate Existing Scripts to the SDK

If you have Python scripts that currently wrap grep or parse file lists, refactor them to use the Cypher Generator (services/llm.py) and the GraphLoader class from codebase_rag/graph_loader.py. This maintains programmatic access while gaining graph traversal capabilities.

Architecture Overview for Migration

Understanding the component hierarchy ensures you map old workflows to the correct new subsystems:

Layer Responsibility Key Implementation
Parser Tree-sitter drives language-agnostic AST extraction. codebase_rag/parsers/ – uses tree_sitter bindings.
Graph Ingestion Translates AST nodes/relationships into Memgraph entities. codebase_rag/graph_updater.py and codebase_rag/graph_loader.py.
Embedding / Semantic Index Generates 768-dim vectors with UniXcoder. codebase_rag/services/embedding.py; configure via CGR_VECTOR_STORE_BACKEND.
Query Engine LLM (via pydantic_ai) converts NL → Cypher. services/cypher_generator.py; orchestrated by codebase_rag/main.py.
CLI / UI Rich-based status bar, prompts, and diff rendering. codebase_rag/constants/cli.py.
Persistence Export / import JSON for offline analysis. export_graph_to_file in codebase_rag/main.py.

All components are wired together by the AppContext singleton (codebase_rag/models.py), which propagates configuration (settings) and runtime state throughout the application.

Code Examples: From Legacy Tools to Code-Graph-RAG

Convert legacy text searches to embedding-based retrieval:

from cgr import embed_code, semantic_search

# Generate embedding for a code snippet (usually done once during indexing)

snippet = """
def validate_user(token: str) -> bool:
    # verify JWT token

    ...
"""
embedding = embed_code(snippet)          # uses UniXcoder under the hood

print(f"Embedding size: {len(embedding)}")   # → 768

# Perform semantic search

results = semantic_search("jwt validation function")
for node_id, score in results:
    print(f"Node {node_id} – similarity {score:.2f}")

The embed_code function resides in codebase_rag/services/embedding.py, while semantic_search reaches the vector store configured via environment variables.

Converting Text Searches to Cypher Queries

Migrate shell scripts that invoke grep to Python using the Graph SDK:

from cgr import GraphLoader, connect_memgraph

# Load an existing exported graph

loader = GraphLoader("my_project_graph.json")
loader.load()

# Old grep:   grep -R "def .*parse.*" .

# New Cypher:

query = """
MATCH (f:Function)
WHERE f.name =~ '(?i).*parse.*'
RETURN f.name AS function, f.file AS file
ORDER BY f.name
"""
results = loader.graph_service.run_cypher(query)
for row in results:
    print(f"{row['function']}  –  {row['file']}")

The GraphLoader class in codebase_rag/graph_loader.py handles JSON deserialization, while graph_service wraps the Memgraph client for Cypher execution.

Automating Code Edits via RAG

Use the interactive CLI for AI-guided refactoring:

cgr chat
> replace the logging call in src/utils.py with logger.info(...)

The system executes three steps defined in codebase_rag/main.py:

  1. Generate a diff via _print_unified_diff showing the proposed change.
  2. Prompt for approval using _process_tool_approvals (referencing UI_TOOL_APPROVAL from codebase_rag/constants/cli.py).
  3. Apply the edit via FileEditor (tools/file_editor.py).

Key Files and Source References

File Role
codebase_rag/main.py Central CLI entry point and interactive loops
codebase_rag/graph_loader.py JSON graph loading and export utilities
codebase_rag/constants/cli.py UI constants, approval prompts, and status messages
codebase_rag/parsers/ Tree-sitter AST extraction and multi-language parsing
codebase_rag/services/embedding.py UniXcoder embedding generation
services/cypher_generator.py LLM-powered natural language to Cypher translation
docs/architecture/overview.md High-level architecture documentation
docs/sdk/semantic-search.md Vector store configuration and semantic search usage

Summary

  • Replace static indexing with a one-time cgr start command to build a language-aware graph in Memgraph.
  • Query the graph using Cypher for exact structural matches or semantic search (UniXcoder embeddings) for conceptual queries.
  • Use the Python SDK (GraphLoader, semantic_search) to refactor existing grep-based scripts.
  • Leverage cgr chat for interactive, AI-assisted code editing with approval workflows.
  • Export graphs to JSON via export_graph_to_file for backup and CI pipeline integration.

Frequently Asked Questions

How do I migrate existing grep-based scripts to Code-Graph-RAG?

Wrap your existing logic using the GraphLoader class from codebase_rag/graph_loader.py and the Cypher execution methods in services/graph_service.py. Replace regex patterns with MATCH clauses that query function names, file paths, and relationships. This preserves your automation while adding cross-file context.

Can I migrate incrementally without disrupting existing workflows?

Yes. Code-Graph-RAG can coexist with traditional tools during the transition. Run cgr start to build the initial graph, then gradually migrate specific queries to the new system while keeping legacy scripts operational. The JSON export feature allows you to maintain offline snapshots without database dependency.

Which programming languages are supported when migrating my codebase?

The system uses Tree-sitter for parsing, supporting any language with a Tree-sitter grammar. The parser implementation in codebase_rag/parsers/ extracts symbols and relationships uniformly across languages, making the migration path consistent regardless of whether you are searching Python, JavaScript, Go, or Rust codebases.

How is cross-file call graph analysis handled during migration?

Traditional tools require manual navigation between files. After migration, the graph schema directly stores CALLS relationships between functions across files. Query these using Cypher: MATCH (caller:Function)-[:CALLS]->(callee:Function) RETURN caller, callee. This is powered by the graph ingestion layer in codebase_rag/graph_updater.py and the AST pattern matching defined in codebase_rag/parsers/ast_grep_patterns/.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →