Architecture of the Semantic Search Feature in Code-Graph-RAG: A Deep Dive into the Pipeline

The semantic search feature in Code-Graph-RAG implements a four-stage pipeline that converts natural language queries into dense vector embeddings, retrieves nearest neighbors from a vector store, hydrates results with metadata from Neo4j via Cypher queries, and exposes the entire flow through agent-compatible Tool wrappers.

Understanding the architecture of the semantic search feature in Code-Graph-RAG reveals how the system bridges free-text questions and structured code repositories. Rather than relying on simple keyword matching, the pipeline leverages dense vector similarity against pre-computed embeddings of code entities stored in a graph database. This design separates concerns between embedding generation, vector retrieval, and graph traversal to deliver fast, contextually relevant results while maintaining agent-friendly APIs.

Core Components of the Semantic Search Pipeline

The system isolates responsibilities across four primary modules, each defined in specific source files within the vitali87/code-graph-rag repository:

Component Role Source File
Embedding Engine Converts text queries into dense vectors using the project's configured embedder. codebase_rag/embedder.py
Vector Store Maintains pre-computed embeddings for all indexed code nodes and executes k-nearest-neighbors lookups. codebase_rag/vector_store.py
Cypher Query Engine Retrieves full node metadata (qualified names, types, file locations) from Neo4j using batched node ID lookups. codebase_rag/cypher_queries.py
Tool Wrapper Exposes the pipeline as asynchronous Tools compatible with LLM agent frameworks and provides source-code extraction utilities. codebase_rag/tools/semantic_search.py

How the Semantic Search Pipeline Works

The semantic_code_search function in codebase_rag/tools/semantic_search.py orchestrates the end-to-end flow through six discrete stages:

Dependency Validation

Before processing, the system verifies that optional semantic search dependencies—such as sentence-transformers and faiss—are installed via has_semantic_dependencies. If these are missing, the function aborts early with a logged warning rather than raising an error.

Query Embedding Generation

The pipeline calls embed_code from codebase_rag/embedder.py to transform the user's natural language query into a dense vector:

from ..embedder import embed_code
query_embedding = embed_code(query)

This embedding aligns the query into the same vector space as the pre-indexed code entities.

The query vector is passed to search_embeddings in codebase_rag/vector_store.py to perform a nearest-neighbor lookup:

from ..vector_store import search_embeddings
search_results = search_embeddings(query_embedding, top_k=top_k, project=project)

This returns a sorted list of (node_id, similarity_score) tuples representing the most relevant code entities.

Graph Metadata Retrieval via Cypher

Rather than hitting the database repeatedly, the system batches node IDs into a single Cypher query using build_nodes_by_ids_query (defined in codebase_rag/cypher_queries.py). This executes one round-trip to Neo4j to fetch complete metadata—including qualified_name, name, type, and file location—for all candidate nodes simultaneously.

Result Formatting and Return

The raw graph results are mapped to SemanticSearchResult objects (defined in codebase_rag/types_defs.py), which attach the similarity scores to the node metadata:

SemanticSearchResult(
    node_id=node_id,
    qualified_name=metadata["qualified_name"],
    name=metadata["name"],
    type=metadata["type"],
    score=round(similarity, 4)
)

The function returns a list of these typed objects, which the surrounding Tool wrapper then renders into human-readable text for agent consumption.

Source Code Retrieval for Agent Workflows

Beyond basic search, the architecture exposes a secondary Tool for retrieving original source code. The get_function_source_code helper in semantic_search.py queries the graph using CYPHER_GET_FUNCTION_SOURCE_LOCATION to obtain file paths and line ranges. After validating the location via validate_source_location, it extracts exact source lines using extract_source_lines from codebase_rag/utils/source_extraction.py.

This is wrapped as create_get_function_source_tool, allowing LLM agents to move from semantic results ("find the JSON parser") to concrete implementation details ("show me the code") in a single conversation turn.

Direct Python API

For programmatic access without agent overhead, import semantic_code_search directly:

from codebase_rag.tools.semantic_search import semantic_code_search

results = semantic_code_search(
    ingestor=ingestor,  # QueryProtocol implementation for Neo4j

    query="read a CSV file into a pandas DataFrame",
    top_k=5,
    project=None,
)

for r in results:
    print(f"{r.qualified_name} (type={r.type}, score={r.score})")

Agent-Compatible Tool Interface

To expose search capabilities to LLM agents, use the factory function:

from codebase_rag.tools.semantic_search import create_semantic_search_tool

semantic_tool = create_semantic_search_tool(ingestor)

response = await semantic_tool.run(
    query="how to parse JSON in Rust",
    top_k=3,
    project="my-rust-project"
)
print(response)  # Formatted text suitable for LLM context windows

Fetching Original Source Code

After obtaining a node_id from search results, retrieve the actual implementation:

from codebase_rag.tools.semantic_search import create_get_function_source_tool

source_tool = create_get_function_source_tool(ingestor)
source = await source_tool.run(node_id=12345)
print(source)  # Formatted source snippet with line numbers

Design Rationale and Performance Characteristics

The architecture prioritizes separation of concerns between embedding logic, vector storage, and graph persistence. This modularity allows independent scaling—such as swapping the FAISS-based vector_store.py implementation for a managed ANN service—without touching the query pipeline or Neo4j schema.

Single-round-trip graph access minimizes database latency by batching node metadata requests into one Cypher query rather than N individual lookups. This proves critical when top_k exceeds single-digit values.

The system implements graceful degradation: missing optional dependencies trigger early returns with logged warnings, and empty vector search results propagate as empty lists rather than exceptions, ensuring LLM agents receive predictable responses even in degraded states.

Summary

  • Four-stage pipeline: Embedding generation → Vector similarity search → Cypher metadata hydration → Tool-wrapped results.
  • Key source files: semantic_search.py orchestrates the flow, embedder.py handles vectorization, vector_store.py manages the ANN index, and cypher_queries.py defines the graph access patterns.
  • Performance optimized: Batched Cypher queries reduce Neo4j round-trips; search_embeddings uses efficient k-NN lookups.
  • Agent-ready: Both create_semantic_search_tool and create_get_function_source_tool provide async interfaces compatible with LLM agent frameworks.
  • Resilient design: Dependency checks and empty-result handling prevent pipeline failures in production environments.

Frequently Asked Questions

The system uses the embed_code function from codebase_rag/embedder.py, which typically supports OpenAI embeddings or local Sentence-Transformers models depending on the project configuration. The vector store expects dense vectors regardless of the specific provider, allowing flexibility in model selection.

How does the semantic search handle projects with millions of code entities?

The architecture delegates scalability to the underlying vector store implementation in codebase_rag/vector_store.py. By default, it uses FAISS for efficient approximate nearest neighbor search, which handles millions of vectors in memory. For larger deployments, the module can be extended to use distributed vector databases without modifying the semantic_code_search logic.

Why does the pipeline separate vector search from Cypher graph queries?

This separation minimizes database load and latency. The vector store (often FAISS) performs the heavy similarity computation, returning only the top-k node IDs. The subsequent Cypher query fetches rich metadata for just those IDs in a single round-trip, avoiding expensive graph traversals during the similarity computation phase.

Can I use semantic search without installing the heavy ML dependencies?

Yes, but with limitations. The has_semantic_dependencies check in codebase_rag/utils/dependencies.py guards against import errors. If sentence-transformers or faiss are missing, the semantic_code_search function logs a warning and returns an empty result list, allowing the broader application to function without crashing while disabling only the semantic search capability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →