How to Use Semantic Search with Vector Embeddings in code‑graph‑rag

Semantic search with vector embeddings converts natural‑language queries into dense vector representations, retrieves similar code entities from a vector store, and fetches matching nodes from the graph database.

The code‑graph‑rag repository implements this pipeline through three integrated components: query embedding, vector store lookup, and graph retrieval. This guide walks through the architecture, implementation files, and practical code examples for leveraging semantic search in your codebase exploration.

How Semantic Search Works in code‑graph‑rag

The semantic search pipeline follows a three‑stage workflow that bridges natural language and structured code graphs.

Step 1: Query Embedding

The process starts in codebase_rag/embedder.py, where the embed_code function transforms your text query into a numerical vector. This function supports two backends:

  • OpenAI embedding API – for production deployments with API access
  • Bundled UnixCoder model – for local, offline operation

The embedding function handles the heavy lifting of converting semantic meaning into a searchable vector space.

Step 2: Vector Store Lookup

With the query vector in hand, search_embeddings in codebase_rag/vector_store.py queries the configured vector store—typically Milvus Lite or a compatible backend. This returns the top‑k matching node IDs paired with similarity scores, ranking results by semantic relevance rather than lexical matching.

Step 3: Graph Retrieval and Result Formatting

The orchestration happens in semantic_code_search within codebase_rag/tools/semantic_search.py. This function:

  1. Validates that semantic dependencies are available via has_semantic_dependencies
  2. Calls the embedder and vector store in sequence
  3. Executes Cypher queries to fetch metadata (qualified name, type, etc.)
  4. Returns structured SemanticSearchResult objects

If dependencies are missing, the function logs a warning and returns an empty list—keeping the system functional without breaking non‑semantic workflows.

Core Data Structure: SemanticSearchResult

Results are encapsulated in the SemanticSearchResult dataclass defined in codebase_rag/types_defs.py:

node_id: int          # Internal graph identifier

qualified_name: str   # Fully qualified symbol name

name: str             # Short symbol name

type: str             # Entity type (function, class, method, etc.)

score: float          # Similarity score (0.0–1.0)

This structure lets downstream consumers work with typed objects rather than raw tuples or dictionaries.

Implementation Example: Direct API Usage

For programmatic access to semantic search, import and call semantic_code_search directly:

from codebase_rag.tools.semantic_search import semantic_code_search
from codebase_rag.services import QueryProtocol  # Neo4j/GraphQL ingestor

# Obtain a configured ingestor for your graph database

ingestor: QueryProtocol = ...  

results = semantic_code_search(
    ingestor,
    query="find all functions that parse JSON",
    top_k=5,
    project="my_project",  # Optional: restrict to specific project

)

for r in results:
    print(f"{r.qualified_name} (type={r.type}) – score {r.score:.3f}")

This low‑level approach gives you full control over result processing and integration with existing pipelines.

Implementation Example: LLM Agent Integration

For agentic workflows, use create_semantic_search_tool to get a ready‑made, formatted interface:

from codebase_rag.tools.semantic_search import create_semantic_search_tool

ingestor: QueryProtocol = ...  # Your graph ingestor

semantic_tool = create_semantic_search_tool(ingestor)

async def ask():
    response = await semantic_tool.run(
        "search for functions that open a file", 
        top_k=3
    )
    print(response)

# Running the coroutine produces formatted output:

# 1. my_module.file_io.open_file (type: function, score: 0.912)

# 2. utils.io.read_file (type: function, score: 0.887)

# 3. legacy.files.open_wrapper (type: function, score: 0.854)

The tool handles async execution, result enumeration, and human‑readable formatting automatically.

Key Source Files

File Responsibility
codebase_rag/tools/semantic_search.py Orchestrates the full pipeline: embedding → lookup → formatting
codebase_rag/embedder.py embed_code and batch variants for vector generation
codebase_rag/vector_store.py search_embeddings abstracts vector store operations
codebase_rag/types_defs.py SemanticSearchResult data model
codebase_rag/utils/dependencies.py has_semantic_dependencies guards optional features

Graceful Degradation

The implementation defensively checks for optional dependencies. When has_semantic_dependencies returns False—meaning embedding libraries or vector store clients aren't installed—semantic_code_search logs a warning and returns an empty list. This keeps the broader code‑graph‑rag system operational even in minimal installation environments.

Summary

  • Query embedding happens via embed_code in embedder.py, supporting OpenAI or UnixCoder
  • Vector search is handled by search_embeddings in vector_store.py against Milvus Lite
  • Result assembly occurs in semantic_code_search in tools/semantic_search.py
  • Use direct API calls for custom integrations or create_semantic_search_tool for LLM agents
  • The system gracefully degrades when semantic dependencies are unavailable

Frequently Asked Questions

What embedding models does code‑graph‑rag support?

code‑graph‑rag supports two embedding backends: the OpenAI embedding API for cloud‑based operation and UnixCoder, a code‑specific model bundled with the repository for local execution. Both are accessible through the unified embed_code interface in codebase_rag/embedder.py.

Can I use semantic search without installing all dependencies?

No—semantic search requires optional dependencies that are checked via has_semantic_dependencies in codebase_rag/utils/dependencies.py. If these are missing, semantic_code_search returns an empty list with a logged warning, allowing the rest of the system to function normally.

How are similarity scores calculated?

Scores come directly from the vector store's distance metric (typically cosine similarity), normalized to a 0.0–1.0 range. Higher scores indicate closer semantic alignment between your natural‑language query and the indexed code entity.

What is the difference between semantic_code_search and create_semantic_search_tool?

semantic_code_search is the low‑level function returning structured SemanticSearchResult objects for programmatic processing. create_semantic_search_tool wraps this functionality into a formatted, async‑compatible interface suitable for LLM agents, handling string serialization and enumeration automatically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →