How to Perform Semantic Search with Code-Graph-RAG
Semantic search with Code-Graph-RAG captures your codebase as a graph, embeds every function and method using OpenAI or UniXcoder, and retrieves relevant snippets by ranking cosine similarity between natural-language queries and code vectors, with optional reranking based on graph proximity.
Code-Graph-RAG (CGR) is an open-source framework that unifies vector-based retrieval with structural code graph information. According to the vitali87/code-graph-rag source code, the semantic search pipeline operates in four distinct stages: graph capture, snippet extraction, embedding generation, and similarity ranking. This guide explains the exact implementation details, file paths, and API calls required to execute semantic search over your own repositories.
The Semantic Search Pipeline
The end-to-end workflow implemented in evals/semantic_search.py transforms a raw source tree into ranked code recommendations through a deterministic multi-stage process.
1. Capturing the Code Graph
The process begins with _capture in codebase_rag/graph_updater.py, which parses your source tree and persists nodes for every first-party function and method. Each node stores metadata including path, start_line, and end_line values that map embeddings back to specific source locations.
2. Extracting Function Snippets
Once the graph exists, function_snippets() (defined in evals/semantic_search.py at lines 42-63) reads the original files using the node’s stored coordinates. This returns a mapping of qualified names (qn) to raw source strings, ensuring that embeddings operate on actual code text rather than abstract symbols.
3. Embedding Snippets and Queries
The embed_code_batch() function in codebase_rag/embedder.py (lines 22-33 and 80-101) handles vectorization. It supports two backends:
- OpenAI API for cloud-based embeddings
- Local UniXcoder for offline, privacy-preserving inference
Results are cached in EmbeddingCache to avoid recomputation across queries. Both code snippets and natural-language queries pass through this same embedding layer to ensure vector space alignment.
4. Ranking by Cosine Similarity
cgr_semantic_ranking() (implemented in evals/semantic_search.py at lines 66-86) computes dot-product similarity between the query embedding and every function embedding. It returns the top-k qualified names most semantically similar to the input query.
5. Optional Graph-Proximity Reranking
For enhanced accuracy, the baseline ranking can be refined using reranked_semantic_ranking() (lines 102-118). This function accepts proximity_edges() data—extracted from call and override relationships in the captured graph—and boosts snippets that are structurally adjacent to highly-ranked candidates. The weight parameter controls the balance between semantic similarity and graph proximity.
Running Semantic Search from the CLI
The CLI entry point at the bottom of evals/semantic_search.py wires all components together and outputs recall@k and MRR metrics.
python -m codebase_rag.evals.semantic_search \
--target path/to/.cgr \
--project-name my_project \
--top-k 5 \
--retain 3
--targetspecifies the directory containing the captured CGR graph--top-ksets the cutoff for recall@k and Mean Reciprocal Rank calculations--retainkeepstop-k × retaincandidates before applying graph-proximity reranking
Programmatic Usage Examples
Ranking Functions in Python
Import the core ranking logic to embed semantic search directly into your applications:
from pathlib import Path
from codebase_rag.evals.semantic_search import (
cgr_semantic_ranking,
function_snippets,
)
target = Path("/path/to/.cgr")
project = "my_project"
queries = ["How does the cache work?", "Create a new user account"]
top_k = 10
# Get ranked qualified names for each query
ranking = cgr_semantic_ranking(target, project, queries, top_k)
for q, hits in ranking.items():
print(f"Query: {q}")
for rank, qn in enumerate(hits, start=1):
print(f" {rank}. {qn}")
Embedding Code Directly
For custom workflows, use the embedder module to vectorize arbitrary code strings:
from codebase_rag.embedder import embed_code, embed_code_batch
snippet = "def add(a, b): return a + b"
vector = embed_code(snippet) # single call (cached)
vectors = embed_code_batch([snippet, "print('hello')"]) # batch processing
Applying Graph-Based Reranking
Enhance initial rankings using structural relationships from the code graph:
from codebase_rag.evals.semantic_search import proximity_edges, reranked_semantic_ranking
# Extract adjacency from call/override relationships
adjacency = proximity_edges(target, project)
# Rerank with 70% weight on graph proximity
reranked = reranked_semantic_ranking(ranking, adjacency, weight=0.7)
Key Files and Architecture
Understanding the following source files is essential for customizing semantic search behavior:
evals/semantic_search.py– Core implementation containingfunction_snippets(),cgr_semantic_ranking(), andreranked_semantic_ranking()codebase_rag/embedder.py– Embedding abstraction withembed_code_batch(), supporting OpenAI and UniXcoder backends viaEmbeddingCachecodebase_rag/graph_updater.py– Contains_capturefor initial graph construction and node metadata extractioncodebase_rag/constants.py– Defines node labels (FUNCTION,METHOD) used throughout the pipelinecodebase_rag/tools/graph_rerank.py– Implements the graph-proximity algorithm consumed byreranked_semantic_ranking()codebase_rag/config.py– Runtime settings for embedding providers, model selection, and API authentication
Summary
- Capture first: Use
_captureingraph_updater.pyto index your codebase with precise source location metadata - Extract accurately:
function_snippets()retrieves exact code blocks using stored line ranges from the graph - Embed consistently:
embed_code_batch()unifies OpenAI and UniXcoder backends with intelligent caching - Rank by similarity:
cgr_semantic_ranking()retrieves top-k matches using cosine similarity between query and code vectors - Rerank structurally: Optional
reranked_semantic_ranking()leveragesproximity_edges()to boost graph-adjacent candidates
Frequently Asked Questions
What embedding models does Code-Graph-RAG support?
Code-Graph-RAG supports two primary backends as implemented in codebase_rag/embedder.py: the OpenAI API (for cloud-based embeddings) and the local UniXcoder model (for offline inference). The active backend is controlled via codebase_rag/config.py settings, and all embeddings are cached on disk to minimize API costs and computation time.
How does graph-proximity reranking improve search results?
Graph-proximity reranking uses proximity_edges() to extract call and override relationships from the captured graph. When reranked_semantic_ranking() is invoked with a weight parameter, it boosts the scores of functions that are structurally connected to already highly-ranked candidates. This compensates for cases where semantic similarity alone might miss contextually relevant code that is tightly coupled in the architecture.
Can I perform semantic search on a partial codebase?
Yes. The _capture function in codebase_rag/graph_updater.py can target specific subdirectories, and the function_snippets() extractor only processes nodes present in the captured graph. As long as the target directory contains a valid .cgr graph snapshot, cgr_semantic_ranking() will search only the indexed functions regardless of total repository size.
Where is the embedding cache stored?
The EmbeddingCache class in codebase_rag/embedder.py persists vectors to disk automatically. The cache location is determined by your configuration in codebase_rag/config.py, ensuring that repeated queries against the same codebase do not trigger redundant embedding generation for unchanged functions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →