How to Use Semantic Search with Vector Embeddings in code‑graph‑rag
Semantic search with vector embeddings converts natural‑language queries into dense vector representations, retrieves similar code entities from a vector store, and fetches matching nodes from the graph database.
The code‑graph‑rag repository implements this pipeline through three integrated components: query embedding, vector store lookup, and graph retrieval. This guide walks through the architecture, implementation files, and practical code examples for leveraging semantic search in your codebase exploration.
How Semantic Search Works in code‑graph‑rag
The semantic search pipeline follows a three‑stage workflow that bridges natural language and structured code graphs.
Step 1: Query Embedding
The process starts in codebase_rag/embedder.py, where the embed_code function transforms your text query into a numerical vector. This function supports two backends:
- OpenAI embedding API – for production deployments with API access
- Bundled UnixCoder model – for local, offline operation
The embedding function handles the heavy lifting of converting semantic meaning into a searchable vector space.
Step 2: Vector Store Lookup
With the query vector in hand, search_embeddings in codebase_rag/vector_store.py queries the configured vector store—typically Milvus Lite or a compatible backend. This returns the top‑k matching node IDs paired with similarity scores, ranking results by semantic relevance rather than lexical matching.
Step 3: Graph Retrieval and Result Formatting
The orchestration happens in semantic_code_search within codebase_rag/tools/semantic_search.py. This function:
- Validates that semantic dependencies are available via
has_semantic_dependencies - Calls the embedder and vector store in sequence
- Executes Cypher queries to fetch metadata (qualified name, type, etc.)
- Returns structured
SemanticSearchResultobjects
If dependencies are missing, the function logs a warning and returns an empty list—keeping the system functional without breaking non‑semantic workflows.
Core Data Structure: SemanticSearchResult
Results are encapsulated in the SemanticSearchResult dataclass defined in codebase_rag/types_defs.py:
node_id: int # Internal graph identifier
qualified_name: str # Fully qualified symbol name
name: str # Short symbol name
type: str # Entity type (function, class, method, etc.)
score: float # Similarity score (0.0–1.0)
This structure lets downstream consumers work with typed objects rather than raw tuples or dictionaries.
Implementation Example: Direct API Usage
For programmatic access to semantic search, import and call semantic_code_search directly:
from codebase_rag.tools.semantic_search import semantic_code_search
from codebase_rag.services import QueryProtocol # Neo4j/GraphQL ingestor
# Obtain a configured ingestor for your graph database
ingestor: QueryProtocol = ...
results = semantic_code_search(
ingestor,
query="find all functions that parse JSON",
top_k=5,
project="my_project", # Optional: restrict to specific project
)
for r in results:
print(f"{r.qualified_name} (type={r.type}) – score {r.score:.3f}")
This low‑level approach gives you full control over result processing and integration with existing pipelines.
Implementation Example: LLM Agent Integration
For agentic workflows, use create_semantic_search_tool to get a ready‑made, formatted interface:
from codebase_rag.tools.semantic_search import create_semantic_search_tool
ingestor: QueryProtocol = ... # Your graph ingestor
semantic_tool = create_semantic_search_tool(ingestor)
async def ask():
response = await semantic_tool.run(
"search for functions that open a file",
top_k=3
)
print(response)
# Running the coroutine produces formatted output:
# 1. my_module.file_io.open_file (type: function, score: 0.912)
# 2. utils.io.read_file (type: function, score: 0.887)
# 3. legacy.files.open_wrapper (type: function, score: 0.854)
The tool handles async execution, result enumeration, and human‑readable formatting automatically.
Key Source Files
| File | Responsibility |
|---|---|
codebase_rag/tools/semantic_search.py |
Orchestrates the full pipeline: embedding → lookup → formatting |
codebase_rag/embedder.py |
embed_code and batch variants for vector generation |
codebase_rag/vector_store.py |
search_embeddings abstracts vector store operations |
codebase_rag/types_defs.py |
SemanticSearchResult data model |
codebase_rag/utils/dependencies.py |
has_semantic_dependencies guards optional features |
Graceful Degradation
The implementation defensively checks for optional dependencies. When has_semantic_dependencies returns False—meaning embedding libraries or vector store clients aren't installed—semantic_code_search logs a warning and returns an empty list. This keeps the broader code‑graph‑rag system operational even in minimal installation environments.
Summary
- Query embedding happens via
embed_codeinembedder.py, supporting OpenAI or UnixCoder - Vector search is handled by
search_embeddingsinvector_store.pyagainst Milvus Lite - Result assembly occurs in
semantic_code_searchintools/semantic_search.py - Use direct API calls for custom integrations or
create_semantic_search_toolfor LLM agents - The system gracefully degrades when semantic dependencies are unavailable
Frequently Asked Questions
What embedding models does code‑graph‑rag support?
code‑graph‑rag supports two embedding backends: the OpenAI embedding API for cloud‑based operation and UnixCoder, a code‑specific model bundled with the repository for local execution. Both are accessible through the unified embed_code interface in codebase_rag/embedder.py.
Can I use semantic search without installing all dependencies?
No—semantic search requires optional dependencies that are checked via has_semantic_dependencies in codebase_rag/utils/dependencies.py. If these are missing, semantic_code_search returns an empty list with a logged warning, allowing the rest of the system to function normally.
How are similarity scores calculated?
Scores come directly from the vector store's distance metric (typically cosine similarity), normalized to a 0.0–1.0 range. Higher scores indicate closer semantic alignment between your natural‑language query and the indexed code entity.
What is the difference between semantic_code_search and create_semantic_search_tool?
semantic_code_search is the low‑level function returning structured SemanticSearchResult objects for programmatic processing. create_semantic_search_tool wraps this functionality into a formatted, async‑compatible interface suitable for LLM agents, handling string serialization and enumeration automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →