How to Use the Python SDK for Graph-Loader, Cypher-Generator, and Semantic-Search

The code-graph-rag Python SDK provides three integrated modules—GraphLoader, CypherGenerator, and semantic_search—that let you parse a codebase into a graph, export it to Neo4j via Cypher, and run natural-language semantic search over the code.

The vitali87/code-graph-rag repository implements a complete retrieval-augmented generation (RAG) pipeline for source code. This article demonstrates how to use the Python SDK for graph-loader, cypher-generator, and semantic-search with practical examples drawn from the actual source implementation.

Understanding the Three Core SDK Components

The SDK is organized around three primary capabilities that work sequentially or independently:

Component Primary Module Key Class / Function
Graph-Loader graph_loader.py GraphLoader — parses source files into an internal AST-backed graph
Cypher-Generator cypher_queries.py CypherGenerator — converts the internal graph to Neo4j Cypher statements
Semantic-Search tools/semantic_search.py semantic_code_search() — embeds and queries code by natural language intent

Each component can be used standalone or orchestrated through the high-level CodebaseRag class in [cgr_state.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cgr_state.py).

Loading a Codebase with GraphLoader

The GraphLoader in [codebase_rag/graph_loader.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) recursively discovers source files, parses them using language-specific visitors registered in [language_spec.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py), and builds an intermediate representation capturing files, classes, functions, variables, and their relationships.

from codebase_rag.graph_loader import GraphLoader

# Initialize loader pointing at your project root

loader = GraphLoader(project_root="/path/to/your/project")

# Build the internal graph representation

graph = loader.load()  # Returns a Graph instance with nodes and edges

The resulting graph object contains:

  • File nodes — source file paths with metadata
  • Symbol nodes — classes, functions, methods, and variables
  • Relationship edges — imports, inheritance, calls, and containment

The loader supports multiple languages through pluggable AST visitors defined in [language_spec.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py).

Generating Cypher with CypherGenerator

The CypherGenerator in [codebase_rag/cypher_queries.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) traverses the internal graph and emits idempotent Cypher MERGE statements suitable for bulk-loading into Neo4j.

from codebase_rag.cypher_queries import CypherGenerator

# Initialize the generator

cypher_gen = CypherGenerator()

# Convert graph to Cypher statements

cypher_statements = cypher_gen.generate_cypher(graph)

# Optional: Execute against Neo4j

# from neo4j import GraphDatabase

# driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

# with driver.session() as session:

#     for statement in cypher_statements:

#         session.run(statement)

The generator creates nodes labeled as :File, :Class, :Function, :Method, and :Variable, with relationships like :DEFINES, :CALLS, :IMPORTS, and :INHERITS_FROM. This schema enables complex graph traversals for code analysis.

The semantic_search module in [codebase_rag/tools/semantic_search.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) provides intent-driven code retrieval using embeddings generated by the [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) module.

from codebase_rag.tools.semantic_search import semantic_code_search

# Search for code matching natural language intent

results = semantic_code_search(
    query="find the function that validates user tokens",
    top_k=5,
    graph=graph,  # Same graph instance from loader

)

# Each result contains source code and metadata

for r in results:
    print(f"File: {r['file_path']}")
    print(f"Function: {r['name']}")
    print(f"Source:\n{r['source_code']}\n")

Behind the scenes, semantic_code_search:

  1. Embeds the query using the configured embedding model
  2. Performs vector similarity search against pre-computed function/method embeddings
  3. Retrieves full source code via get_function_source_code()
  4. Returns ranked results with provenance metadata

Complete SDK Workflow Example

For most use cases, orchestrate all three components through the high-level CodebaseRag class:

from codebase_rag import CodebaseRag

# Initialize the SDK with project path

cgr = CodebaseRag(project_root="my_service")

# Build graph and optionally push to configured Neo4j instance

cgr.build_graph()  # Internally uses GraphLoader + CypherGenerator

# Execute semantic search

snippets = cgr.semantic_search(
    "list all HTTP endpoints that require authentication",
    top_k=3
)

for snippet in snippets:
    print(snippet)

This single entry point handles dependency injection, caching, and configuration management across the three core modules.

Customizing the Embedding Model

To use a custom embedding provider for semantic search, instantiate the Embedder directly and create a LangChain-compatible tool:

from codebase_rag.embedder import Embedder
from codebase_rag.tools.semantic_search import create_semantic_search_tool

# Configure custom embedding model

embedder = Embedder(model_name="sentence-transformers/all-MiniLM-L6-v2")

# Create searchable tool instance

search_tool = create_semantic_search_tool(ingestor=embedder)

# Execute search (LangChain-compatible interface)

result = search_tool.function(
    natural_language_query="extract the password hashing routine",
    top_k=2
)
print(result)

The Embedder class in [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) abstracts model loading, batch encoding, and vector persistence, allowing seamless swapping between OpenAI, HuggingFace, or local embedding models.

Key Source Files Reference

File Purpose Direct Link
graph_loader.py AST parsing and internal graph construction View source
cypher_queries.py Cypher statement generation from internal graph View source
tools/semantic_search.py Natural language search over code embeddings View source
embedder.py Embedding model wrapper and vector operations View source
cgr_state.py High-level SDK orchestrator (CodebaseRag) View source
language_spec.py Language-specific parser registrations View source

Summary

Frequently Asked Questions

How do I install the code-graph-rag Python SDK?

The repository is installable via pip from source. Clone the repository and run pip install -e . from the root directory. Per the repository structure, installable packages are defined under the codebase_rag/ namespace with dependencies specified in pyproject.toml or setup.py.

Can I use the graph-loader without Neo4j?

Yes. The GraphLoader produces an internal graph representation independent of any database. You can analyze, traverse, or serialize this graph without invoking the CypherGenerator. The internal graph format is native Python objects suitable for custom processing.

The Embedder class supports any model compatible with the sentence-transformers library or OpenAI's embedding API. Pass the model identifier to Embedder(model_name="...") or set via environment configuration. The default model and dimensionality are configurable in [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →