How to Use the Python SDK for Graph-Loader, Cypher-Generator, and Semantic-Search
The code-graph-rag Python SDK provides three integrated modules—GraphLoader, CypherGenerator, and semantic_search—that let you parse a codebase into a graph, export it to Neo4j via Cypher, and run natural-language semantic search over the code.
The vitali87/code-graph-rag repository implements a complete retrieval-augmented generation (RAG) pipeline for source code. This article demonstrates how to use the Python SDK for graph-loader, cypher-generator, and semantic-search with practical examples drawn from the actual source implementation.
Understanding the Three Core SDK Components
The SDK is organized around three primary capabilities that work sequentially or independently:
| Component | Primary Module | Key Class / Function |
|---|---|---|
| Graph-Loader | graph_loader.py |
GraphLoader — parses source files into an internal AST-backed graph |
| Cypher-Generator | cypher_queries.py |
CypherGenerator — converts the internal graph to Neo4j Cypher statements |
| Semantic-Search | tools/semantic_search.py |
semantic_code_search() — embeds and queries code by natural language intent |
Each component can be used standalone or orchestrated through the high-level CodebaseRag class in [cgr_state.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cgr_state.py).
Loading a Codebase with GraphLoader
The GraphLoader in [codebase_rag/graph_loader.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) recursively discovers source files, parses them using language-specific visitors registered in [language_spec.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py), and builds an intermediate representation capturing files, classes, functions, variables, and their relationships.
from codebase_rag.graph_loader import GraphLoader
# Initialize loader pointing at your project root
loader = GraphLoader(project_root="/path/to/your/project")
# Build the internal graph representation
graph = loader.load() # Returns a Graph instance with nodes and edges
The resulting graph object contains:
- File nodes — source file paths with metadata
- Symbol nodes — classes, functions, methods, and variables
- Relationship edges — imports, inheritance, calls, and containment
The loader supports multiple languages through pluggable AST visitors defined in [language_spec.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/language_spec.py).
Generating Cypher with CypherGenerator
The CypherGenerator in [codebase_rag/cypher_queries.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) traverses the internal graph and emits idempotent Cypher MERGE statements suitable for bulk-loading into Neo4j.
from codebase_rag.cypher_queries import CypherGenerator
# Initialize the generator
cypher_gen = CypherGenerator()
# Convert graph to Cypher statements
cypher_statements = cypher_gen.generate_cypher(graph)
# Optional: Execute against Neo4j
# from neo4j import GraphDatabase
# driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))
# with driver.session() as session:
# for statement in cypher_statements:
# session.run(statement)
The generator creates nodes labeled as :File, :Class, :Function, :Method, and :Variable, with relationships like :DEFINES, :CALLS, :IMPORTS, and :INHERITS_FROM. This schema enables complex graph traversals for code analysis.
Running Semantic Search
The semantic_search module in [codebase_rag/tools/semantic_search.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) provides intent-driven code retrieval using embeddings generated by the [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) module.
from codebase_rag.tools.semantic_search import semantic_code_search
# Search for code matching natural language intent
results = semantic_code_search(
query="find the function that validates user tokens",
top_k=5,
graph=graph, # Same graph instance from loader
)
# Each result contains source code and metadata
for r in results:
print(f"File: {r['file_path']}")
print(f"Function: {r['name']}")
print(f"Source:\n{r['source_code']}\n")
Behind the scenes, semantic_code_search:
- Embeds the query using the configured embedding model
- Performs vector similarity search against pre-computed function/method embeddings
- Retrieves full source code via
get_function_source_code() - Returns ranked results with provenance metadata
Complete SDK Workflow Example
For most use cases, orchestrate all three components through the high-level CodebaseRag class:
from codebase_rag import CodebaseRag
# Initialize the SDK with project path
cgr = CodebaseRag(project_root="my_service")
# Build graph and optionally push to configured Neo4j instance
cgr.build_graph() # Internally uses GraphLoader + CypherGenerator
# Execute semantic search
snippets = cgr.semantic_search(
"list all HTTP endpoints that require authentication",
top_k=3
)
for snippet in snippets:
print(snippet)
This single entry point handles dependency injection, caching, and configuration management across the three core modules.
Customizing the Embedding Model
To use a custom embedding provider for semantic search, instantiate the Embedder directly and create a LangChain-compatible tool:
from codebase_rag.embedder import Embedder
from codebase_rag.tools.semantic_search import create_semantic_search_tool
# Configure custom embedding model
embedder = Embedder(model_name="sentence-transformers/all-MiniLM-L6-v2")
# Create searchable tool instance
search_tool = create_semantic_search_tool(ingestor=embedder)
# Execute search (LangChain-compatible interface)
result = search_tool.function(
natural_language_query="extract the password hashing routine",
top_k=2
)
print(result)
The Embedder class in [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) abstracts model loading, batch encoding, and vector persistence, allowing seamless swapping between OpenAI, HuggingFace, or local embedding models.
Key Source Files Reference
| File | Purpose | Direct Link |
|---|---|---|
graph_loader.py |
AST parsing and internal graph construction | View source |
cypher_queries.py |
Cypher statement generation from internal graph | View source |
tools/semantic_search.py |
Natural language search over code embeddings | View source |
embedder.py |
Embedding model wrapper and vector operations | View source |
cgr_state.py |
High-level SDK orchestrator (CodebaseRag) |
View source |
language_spec.py |
Language-specific parser registrations | View source |
Summary
GraphLoaderin [graph_loader.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) ingests source code across multiple languages into a unified graph representationCypherGeneratorin [cypher_queries.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) exports this graph to Neo4j-compatible Cypher statementssemantic_code_searchin [tools/semantic_search.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) enables natural language code retrieval using embeddings- The
CodebaseRagclass in [cgr_state.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cgr_state.py) provides a unified SDK interface for all three capabilities - All components support customization: embedding models, language parsers, and Neo4j connection parameters
Frequently Asked Questions
How do I install the code-graph-rag Python SDK?
The repository is installable via pip from source. Clone the repository and run pip install -e . from the root directory. Per the repository structure, installable packages are defined under the codebase_rag/ namespace with dependencies specified in pyproject.toml or setup.py.
Can I use the graph-loader without Neo4j?
Yes. The GraphLoader produces an internal graph representation independent of any database. You can analyze, traverse, or serialize this graph without invoking the CypherGenerator. The internal graph format is native Python objects suitable for custom processing.
What embedding models work with semantic-search?
The Embedder class supports any model compatible with the sentence-transformers library or OpenAI's embedding API. Pass the model identifier to Embedder(model_name="...") or set via environment configuration. The default model and dimensionality are configurable in [embedder.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →