# Architecture of the Semantic Search Feature in Code-Graph-RAG: A Deep Dive into the Pipeline

> Explore the four-stage pipeline architecture of Code-Graph-RAG's semantic search. Learn how it converts queries to vectors, retrieves data, and hydrates results for advanced code understanding. Dive deep into the pipeline!

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: architecture
- Published: 2026-08-18

---

**The semantic search feature in Code-Graph-RAG implements a four-stage pipeline that converts natural language queries into dense vector embeddings, retrieves nearest neighbors from a vector store, hydrates results with metadata from Neo4j via Cypher queries, and exposes the entire flow through agent-compatible Tool wrappers.**

Understanding the **architecture of the semantic search feature in Code-Graph-RAG** reveals how the system bridges free-text questions and structured code repositories. Rather than relying on simple keyword matching, the pipeline leverages dense vector similarity against pre-computed embeddings of code entities stored in a graph database. This design separates concerns between embedding generation, vector retrieval, and graph traversal to deliver fast, contextually relevant results while maintaining agent-friendly APIs.

## Core Components of the Semantic Search Pipeline

The system isolates responsibilities across four primary modules, each defined in specific source files within the `vitali87/code-graph-rag` repository:

| Component | Role | Source File |
|-----------|------|-------------|
| **Embedding Engine** | Converts text queries into dense vectors using the project's configured embedder. | [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) |
| **Vector Store** | Maintains pre-computed embeddings for all indexed code nodes and executes k-nearest-neighbors lookups. | [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) |
| **Cypher Query Engine** | Retrieves full node metadata (qualified names, types, file locations) from Neo4j using batched node ID lookups. | [`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) |
| **Tool Wrapper** | Exposes the pipeline as asynchronous Tools compatible with LLM agent frameworks and provides source-code extraction utilities. | [`codebase_rag/tools/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) |

## How the Semantic Search Pipeline Works

The `semantic_code_search` function in [`codebase_rag/tools/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) orchestrates the end-to-end flow through six discrete stages:

### Dependency Validation

Before processing, the system verifies that optional semantic search dependencies—such as `sentence-transformers` and `faiss`—are installed via `has_semantic_dependencies`. If these are missing, the function aborts early with a logged warning rather than raising an error.

### Query Embedding Generation

The pipeline calls `embed_code` from [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) to transform the user's natural language query into a dense vector:

```python
from ..embedder import embed_code
query_embedding = embed_code(query)

```

This embedding aligns the query into the same vector space as the pre-indexed code entities.

### Vector Similarity Search

The query vector is passed to `search_embeddings` in [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) to perform a nearest-neighbor lookup:

```python
from ..vector_store import search_embeddings
search_results = search_embeddings(query_embedding, top_k=top_k, project=project)

```

This returns a sorted list of `(node_id, similarity_score)` tuples representing the most relevant code entities.

### Graph Metadata Retrieval via Cypher

Rather than hitting the database repeatedly, the system batches node IDs into a single Cypher query using `build_nodes_by_ids_query` (defined in [`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py)). This executes one round-trip to Neo4j to fetch complete metadata—including `qualified_name`, `name`, `type`, and file location—for all candidate nodes simultaneously.

### Result Formatting and Return

The raw graph results are mapped to `SemanticSearchResult` objects (defined in [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py)), which attach the similarity scores to the node metadata:

```python
SemanticSearchResult(
    node_id=node_id,
    qualified_name=metadata["qualified_name"],
    name=metadata["name"],
    type=metadata["type"],
    score=round(similarity, 4)
)

```

The function returns a list of these typed objects, which the surrounding Tool wrapper then renders into human-readable text for agent consumption.

## Source Code Retrieval for Agent Workflows

Beyond basic search, the architecture exposes a secondary Tool for retrieving original source code. The `get_function_source_code` helper in [`semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/semantic_search.py) queries the graph using `CYPHER_GET_FUNCTION_SOURCE_LOCATION` to obtain file paths and line ranges. After validating the location via `validate_source_location`, it extracts exact source lines using `extract_source_lines` from [`codebase_rag/utils/source_extraction.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/source_extraction.py).

This is wrapped as `create_get_function_source_tool`, allowing LLM agents to move from semantic results ("find the JSON parser") to concrete implementation details ("show me the code") in a single conversation turn.

## Implementation Example: Running a Semantic Search

### Direct Python API

For programmatic access without agent overhead, import `semantic_code_search` directly:

```python
from codebase_rag.tools.semantic_search import semantic_code_search

results = semantic_code_search(
    ingestor=ingestor,  # QueryProtocol implementation for Neo4j

    query="read a CSV file into a pandas DataFrame",
    top_k=5,
    project=None,
)

for r in results:
    print(f"{r.qualified_name} (type={r.type}, score={r.score})")

```

### Agent-Compatible Tool Interface

To expose search capabilities to LLM agents, use the factory function:

```python
from codebase_rag.tools.semantic_search import create_semantic_search_tool

semantic_tool = create_semantic_search_tool(ingestor)

response = await semantic_tool.run(
    query="how to parse JSON in Rust",
    top_k=3,
    project="my-rust-project"
)
print(response)  # Formatted text suitable for LLM context windows

```

### Fetching Original Source Code

After obtaining a `node_id` from search results, retrieve the actual implementation:

```python
from codebase_rag.tools.semantic_search import create_get_function_source_tool

source_tool = create_get_function_source_tool(ingestor)
source = await source_tool.run(node_id=12345)
print(source)  # Formatted source snippet with line numbers

```

## Design Rationale and Performance Characteristics

The architecture prioritizes **separation of concerns** between embedding logic, vector storage, and graph persistence. This modularity allows independent scaling—such as swapping the FAISS-based [`vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/vector_store.py) implementation for a managed ANN service—without touching the query pipeline or Neo4j schema.

**Single-round-trip graph access** minimizes database latency by batching node metadata requests into one Cypher query rather than N individual lookups. This proves critical when `top_k` exceeds single-digit values.

The system implements **graceful degradation**: missing optional dependencies trigger early returns with logged warnings, and empty vector search results propagate as empty lists rather than exceptions, ensuring LLM agents receive predictable responses even in degraded states.

## Summary

- **Four-stage pipeline**: Embedding generation → Vector similarity search → Cypher metadata hydration → Tool-wrapped results.
- **Key source files**: [`semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/semantic_search.py) orchestrates the flow, [`embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/embedder.py) handles vectorization, [`vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/vector_store.py) manages the ANN index, and [`cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/cypher_queries.py) defines the graph access patterns.
- **Performance optimized**: Batched Cypher queries reduce Neo4j round-trips; `search_embeddings` uses efficient k-NN lookups.
- **Agent-ready**: Both `create_semantic_search_tool` and `create_get_function_source_tool` provide async interfaces compatible with LLM agent frameworks.
- **Resilient design**: Dependency checks and empty-result handling prevent pipeline failures in production environments.

## Frequently Asked Questions

### What embedding models does Code-Graph-RAG support for semantic search?

The system uses the `embed_code` function from [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py), which typically supports OpenAI embeddings or local Sentence-Transformers models depending on the project configuration. The vector store expects dense vectors regardless of the specific provider, allowing flexibility in model selection.

### How does the semantic search handle projects with millions of code entities?

The architecture delegates scalability to the underlying vector store implementation in [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py). By default, it uses FAISS for efficient approximate nearest neighbor search, which handles millions of vectors in memory. For larger deployments, the module can be extended to use distributed vector databases without modifying the `semantic_code_search` logic.

### Why does the pipeline separate vector search from Cypher graph queries?

This separation minimizes database load and latency. The vector store (often FAISS) performs the heavy similarity computation, returning only the top-k node IDs. The subsequent Cypher query fetches rich metadata for just those IDs in a single round-trip, avoiding expensive graph traversals during the similarity computation phase.

### Can I use semantic search without installing the heavy ML dependencies?

Yes, but with limitations. The `has_semantic_dependencies` check in [`codebase_rag/utils/dependencies.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/dependencies.py) guards against import errors. If `sentence-transformers` or `faiss` are missing, the `semantic_code_search` function logs a warning and returns an empty result list, allowing the broader application to function without crashing while disabling only the semantic search capability.