Underlying Technologies in Code-Graph-RAG: Python, Neo4j, and LLM Integration Explained
Code-Graph-RAG leverages Python 3.11, Neo4j for graph storage, Tree-sitter for AST parsing, PyTorch with Hugging Face Transformers for embeddings, FAISS for vector search, and LangChain for LLM orchestration to enable retrieval-augmented generation over code repositories.
The vitali87/code-graph-rag repository implements a comprehensive framework that transforms static codebases into interactive knowledge graphs. By examining the underlying technologies used in code-graph-rag, developers can understand how the system bridges structural code analysis with semantic vector search to power intelligent code assistance.
Core Technology Stack
Programming Language and Infrastructure
The entire framework is built on Python 3.11, with all implementation code residing in the *.py modules under codebase_rag/. The command-line interface relies on Click for argument parsing and user interaction, while Pydantic provides typed configuration management for runtime settings. Progress visualization during long-running operations is handled by Tqdm, ensuring transparent feedback during graph construction and embedding generation.
Graph Database Layer
Neo4j serves as the persistent storage backend for the code knowledge graph. The system utilizes the official neo4j Python driver to establish connections and execute Cypher queries. In codebase_rag/graph_loader.py, the GraphLoader class manages node and edge creation, persisting structural relationships such as function calls, class hierarchies, and import dependencies. This enables complex graph traversals that pure vector search cannot achieve.
Code Parsing and AST Analysis
To extract structural meaning from source files, the framework employs Tree-sitter via the tree_sitter Python bindings. The codebase_rag/parser_loader.py module dynamically selects the appropriate parser based on file extensions, supporting languages including Python, Java, Rust, Scala, PHP, and Go. This Abstract Syntax Tree (AST) representation allows the system to fingerprint code elements and convert them into graph nodes with precise metadata about function boundaries, variable scopes, and control flow.
Vector Embeddings and Neural Networks
The semantic layer relies on PyTorch and the Hugging Face Transformers library to generate high-dimensional vector embeddings. In codebase_rag/embedder.py, the CodeEmbedder class wraps models such as sentence-transformers/all-mpnet-base-v2 to encode source code snippets into dense vectors. These embeddings are stored in FAISS (Facebook AI Similarity Search), which provides efficient nearest-neighbor look-up capabilities for semantic code search. The FAISSStore class abstracts the indexing and retrieval operations, enabling similarity scoring across codebases.
LLM Integration and Orchestration
LangChain provides the abstraction layer for large language model interactions, handling prompt construction, response parsing, and chain composition. This integration allows the RAG pipeline to combine retrieval results from both the Neo4j graph structure and the FAISS vector store, feeding relevant context into LLM prompts for enhanced code generation and explanation tasks.
System Architecture and Key Components
The framework's architecture centers on the CGRState class defined in codebase_rag/cgr_state.py, which maintains runtime state including the graph representation, vector store instance, and configuration settings.
In codebase_rag/main.py, the entry point orchestrates the complete pipeline: initializing state, triggering graph construction via state.build_graph(), managing embedding generation through CodeEmbedder, and persisting results to Neo4j via GraphLoader.persist(). The codebase_rag/utils/dependencies.py module provides helper functions to detect optional dependencies like torch, transformers, and neo4j, ensuring graceful degradation when specific components are unavailable.
For performance evaluation, benchmarks/bench_graph_loader.py contains scripts that measure the throughput of graph loading and Cypher query execution, while examples/graph_export_example.py demonstrates exporting the constructed graph to formats like JSON and CSV.
Practical Implementation Examples
Creating a Graph from a Local Source Tree
The following example demonstrates initializing the framework and persisting a code graph to Neo4j:
from codebase_rag.cgr_state import CGRState
from codebase_rag.graph_loader import GraphLoader
# Initialise state with default config
state = CGRState(root_path="~/my_project")
# Build the code-graph (parses files, creates nodes & edges)
state.build_graph()
# Persist the graph into Neo4j
GraphLoader(state).persist()
Running a Semantic Code Search
This snippet illustrates encoding code snippets and querying the FAISS vector store:
from codebase_rag.embedder import CodeEmbedder
from codebase_rag.vector_store import FAISSStore
embedder = CodeEmbedder(model="sentence-transformers/all-mpnet-base-v2")
faiss = FAISSStore(dim=embedder.dim)
# Index a snippet
vec = embedder.encode("def foo(x): return x*2")
faiss.add(vec, metadata={"file": "utils.py", "line": 10})
# Query the index
query_vec = embedder.encode("how to double a number")
hits = faiss.search(query_vec, k=5)
print(hits) # → list of similar code snippets
Hybrid Structural and Semantic Search via Cypher
Combine graph relationships with embedding similarity using Cypher queries:
from codebase_rag.graph_loader import GraphLoader
g = GraphLoader(state).connect()
# Find all functions that call `parse_json` and are semantically similar
cypher = """
MATCH (f:Function)-[:CALLS]->(c:Function {name: 'parse_json'})
WHERE f.embedding_score >= $sim_threshold
RETURN f.name, f.file, f.line
"""
results = g.run_query(cypher, {"sim_threshold": 0.8})
for r in results:
print(r)
Summary
- Code-Graph-RAG integrates Neo4j for structural graph storage and FAISS for semantic vector search to create a hybrid retrieval system.
- Tree-sitter parsers in
codebase_rag/parser_loader.pyenable multi-language AST analysis without custom parsers for each language. - The embedding pipeline in
codebase_rag/embedder.pyuses PyTorch and Transformers to generate code vectors, storing them efficiently via FAISS. - LangChain orchestrates the RAG workflow, combining Cypher graph queries with vector similarity search for enhanced LLM context.
- The modular architecture separates concerns between state management (
cgr_state.py), graph persistence (graph_loader.py), and CLI interaction (main.py).
Frequently Asked Questions
What programming languages does Code-Graph-RAG support for parsing?
Code-Graph-RAG supports multiple languages including Python, Java, Rust, Scala, PHP, and Go through Tree-sitter parsers. The codebase_rag/parser_loader.py module dynamically selects the appropriate grammar based on file extensions, enabling consistent AST extraction across different codebases without requiring language-specific tooling.
How does the system store and query vector embeddings?
The framework stores embeddings in FAISS (Facebook AI Similarity Search), a library optimized for fast nearest-neighbor search in high-dimensional spaces. The FAISSStore class handles indexing and retrieval, while codebase_rag/embedder.py manages the PyTorch-based encoding using models like all-mpnet-base-v2 from the Transformers library.
What is the role of Neo4j in the Code-Graph-RAG architecture?
Neo4j acts as the persistent graph database that stores structural relationships extracted from source code, such as function calls, class hierarchies, and import dependencies. The GraphLoader class in codebase_rag/graph_loader.py manages database connections and executes Cypher queries, enabling complex graph traversals that complement semantic vector search for comprehensive code retrieval.
How does the RAG pipeline combine graph structure with vector similarity?
The system uses LangChain to orchestrate a hybrid retrieval process. It first queries the Neo4j graph using Cypher to identify structurally relevant code entities (e.g., functions calling a specific method), then filters or ranks these results using FAISS vector similarity scores. This approach, implemented across graph_loader.py and the embedding modules, ensures retrieved context is both semantically similar and structurally accurate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →