# Underlying Technologies in Code-Graph-RAG: Python, Neo4j, and LLM Integration Explained

> Discover the core technologies powering Code-Graph-RAG. Explore Python, Neo4j, Tree-sitter, PyTorch, FAISS, and LangChain for efficient code retrieval and generation.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: deep-dive
- Published: 2026-08-18

---

**Code-Graph-RAG leverages Python 3.11, Neo4j for graph storage, Tree-sitter for AST parsing, PyTorch with Hugging Face Transformers for embeddings, FAISS for vector search, and LangChain for LLM orchestration to enable retrieval-augmented generation over code repositories.**

The vitali87/code-graph-rag repository implements a comprehensive framework that transforms static codebases into interactive knowledge graphs. By examining the underlying technologies used in code-graph-rag, developers can understand how the system bridges structural code analysis with semantic vector search to power intelligent code assistance.

## Core Technology Stack

### Programming Language and Infrastructure

The entire framework is built on **Python 3.11**, with all implementation code residing in the `*.py` modules under `codebase_rag/`. The command-line interface relies on **Click** for argument parsing and user interaction, while **Pydantic** provides typed configuration management for runtime settings. Progress visualization during long-running operations is handled by **Tqdm**, ensuring transparent feedback during graph construction and embedding generation.

### Graph Database Layer

**Neo4j** serves as the persistent storage backend for the code knowledge graph. The system utilizes the official `neo4j` Python driver to establish connections and execute **Cypher** queries. In [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py), the `GraphLoader` class manages node and edge creation, persisting structural relationships such as function calls, class hierarchies, and import dependencies. This enables complex graph traversals that pure vector search cannot achieve.

### Code Parsing and AST Analysis

To extract structural meaning from source files, the framework employs **Tree-sitter** via the `tree_sitter` Python bindings. The [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) module dynamically selects the appropriate parser based on file extensions, supporting languages including Python, Java, Rust, Scala, PHP, and Go. This Abstract Syntax Tree (AST) representation allows the system to fingerprint code elements and convert them into graph nodes with precise metadata about function boundaries, variable scopes, and control flow.

### Vector Embeddings and Neural Networks

The semantic layer relies on **PyTorch** and the **Hugging Face Transformers** library to generate high-dimensional vector embeddings. In [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py), the `CodeEmbedder` class wraps models such as `sentence-transformers/all-mpnet-base-v2` to encode source code snippets into dense vectors. These embeddings are stored in **FAISS** (Facebook AI Similarity Search), which provides efficient nearest-neighbor look-up capabilities for semantic code search. The `FAISSStore` class abstracts the indexing and retrieval operations, enabling similarity scoring across codebases.

### LLM Integration and Orchestration

**LangChain** provides the abstraction layer for large language model interactions, handling prompt construction, response parsing, and chain composition. This integration allows the RAG pipeline to combine retrieval results from both the Neo4j graph structure and the FAISS vector store, feeding relevant context into LLM prompts for enhanced code generation and explanation tasks.

## System Architecture and Key Components

The framework's architecture centers on the `CGRState` class defined in [`codebase_rag/cgr_state.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cgr_state.py), which maintains runtime state including the graph representation, vector store instance, and configuration settings.

In [`codebase_rag/main.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/main.py), the entry point orchestrates the complete pipeline: initializing state, triggering graph construction via `state.build_graph()`, managing embedding generation through `CodeEmbedder`, and persisting results to Neo4j via `GraphLoader.persist()`. The [`codebase_rag/utils/dependencies.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/dependencies.py) module provides helper functions to detect optional dependencies like `torch`, `transformers`, and `neo4j`, ensuring graceful degradation when specific components are unavailable.

For performance evaluation, [`benchmarks/bench_graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/benchmarks/bench_graph_loader.py) contains scripts that measure the throughput of graph loading and Cypher query execution, while [`examples/graph_export_example.py`](https://github.com/vitali87/code-graph-rag/blob/main/examples/graph_export_example.py) demonstrates exporting the constructed graph to formats like JSON and CSV.

## Practical Implementation Examples

### Creating a Graph from a Local Source Tree

The following example demonstrates initializing the framework and persisting a code graph to Neo4j:

```python
from codebase_rag.cgr_state import CGRState
from codebase_rag.graph_loader import GraphLoader

# Initialise state with default config

state = CGRState(root_path="~/my_project")

# Build the code-graph (parses files, creates nodes & edges)

state.build_graph()

# Persist the graph into Neo4j

GraphLoader(state).persist()

```

### Running a Semantic Code Search

This snippet illustrates encoding code snippets and querying the FAISS vector store:

```python
from codebase_rag.embedder import CodeEmbedder
from codebase_rag.vector_store import FAISSStore

embedder = CodeEmbedder(model="sentence-transformers/all-mpnet-base-v2")
faiss = FAISSStore(dim=embedder.dim)

# Index a snippet

vec = embedder.encode("def foo(x): return x*2")
faiss.add(vec, metadata={"file": "utils.py", "line": 10})

# Query the index

query_vec = embedder.encode("how to double a number")
hits = faiss.search(query_vec, k=5)
print(hits)   # → list of similar code snippets

```

### Hybrid Structural and Semantic Search via Cypher

Combine graph relationships with embedding similarity using Cypher queries:

```python
from codebase_rag.graph_loader import GraphLoader

g = GraphLoader(state).connect()

# Find all functions that call `parse_json` and are semantically similar

cypher = """
MATCH (f:Function)-[:CALLS]->(c:Function {name: 'parse_json'})
WHERE f.embedding_score >= $sim_threshold
RETURN f.name, f.file, f.line
"""

results = g.run_query(cypher, {"sim_threshold": 0.8})
for r in results:
    print(r)

```

## Summary

- **Code-Graph-RAG** integrates **Neo4j** for structural graph storage and **FAISS** for semantic vector search to create a hybrid retrieval system.
- **Tree-sitter** parsers in [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) enable multi-language AST analysis without custom parsers for each language.
- The embedding pipeline in [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) uses **PyTorch** and **Transformers** to generate code vectors, storing them efficiently via **FAISS**.
- **LangChain** orchestrates the RAG workflow, combining Cypher graph queries with vector similarity search for enhanced LLM context.
- The modular architecture separates concerns between state management ([`cgr_state.py`](https://github.com/vitali87/code-graph-rag/blob/main/cgr_state.py)), graph persistence ([`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py)), and CLI interaction ([`main.py`](https://github.com/vitali87/code-graph-rag/blob/main/main.py)).

## Frequently Asked Questions

### What programming languages does Code-Graph-RAG support for parsing?

Code-Graph-RAG supports multiple languages including Python, Java, Rust, Scala, PHP, and Go through **Tree-sitter** parsers. The [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) module dynamically selects the appropriate grammar based on file extensions, enabling consistent AST extraction across different codebases without requiring language-specific tooling.

### How does the system store and query vector embeddings?

The framework stores embeddings in **FAISS** (Facebook AI Similarity Search), a library optimized for fast nearest-neighbor search in high-dimensional spaces. The `FAISSStore` class handles indexing and retrieval, while [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) manages the **PyTorch**-based encoding using models like `all-mpnet-base-v2` from the **Transformers** library.

### What is the role of Neo4j in the Code-Graph-RAG architecture?

**Neo4j** acts as the persistent graph database that stores structural relationships extracted from source code, such as function calls, class hierarchies, and import dependencies. The `GraphLoader` class in [`codebase_rag/graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_loader.py) manages database connections and executes **Cypher** queries, enabling complex graph traversals that complement semantic vector search for comprehensive code retrieval.

### How does the RAG pipeline combine graph structure with vector similarity?

The system uses **LangChain** to orchestrate a hybrid retrieval process. It first queries the **Neo4j** graph using Cypher to identify structurally relevant code entities (e.g., functions calling a specific method), then filters or ranks these results using **FAISS** vector similarity scores. This approach, implemented across [`graph_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_loader.py) and the embedding modules, ensures retrieved context is both semantically similar and structurally accurate.