# How to Use Semantic Search with Vector Embeddings in code‑graph‑rag

> Learn how to implement semantic search with vector embeddings in code-graph-rag. Discover how to convert queries, retrieve code, and fetch graph nodes for powerful code understanding.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-20

---

**Semantic search with vector embeddings converts natural‑language queries into dense vector representations, retrieves similar code entities from a vector store, and fetches matching nodes from the graph database.**

The **code‑graph‑rag** repository implements this pipeline through three integrated components: query embedding, vector store lookup, and graph retrieval. This guide walks through the architecture, implementation files, and practical code examples for leveraging semantic search in your codebase exploration.

## How Semantic Search Works in code‑graph‑rag

The semantic search pipeline follows a three‑stage workflow that bridges natural language and structured code graphs.

### Step 1: Query Embedding

The process starts in [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py), where the `embed_code` function transforms your text query into a numerical vector. This function supports two backends:

- **OpenAI embedding API** – for production deployments with API access
- **Bundled UnixCoder model** – for local, offline operation

The embedding function handles the heavy lifting of converting semantic meaning into a searchable vector space.

### Step 2: Vector Store Lookup

With the query vector in hand, `search_embeddings` in [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) queries the configured vector store—typically **Milvus Lite** or a compatible backend. This returns the top‑k matching node IDs paired with **similarity scores**, ranking results by semantic relevance rather than lexical matching.

### Step 3: Graph Retrieval and Result Formatting

The orchestration happens in `semantic_code_search` within [`codebase_rag/tools/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py). This function:

1. Validates that semantic dependencies are available via `has_semantic_dependencies`
2. Calls the embedder and vector store in sequence
3. Executes Cypher queries to fetch metadata (qualified name, type, etc.)
4. Returns structured `SemanticSearchResult` objects

If dependencies are missing, the function logs a warning and returns an empty list—keeping the system functional without breaking non‑semantic workflows.

## Core Data Structure: SemanticSearchResult

Results are encapsulated in the `SemanticSearchResult` dataclass defined in [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py):

```python
node_id: int          # Internal graph identifier

qualified_name: str   # Fully qualified symbol name

name: str             # Short symbol name

type: str             # Entity type (function, class, method, etc.)

score: float          # Similarity score (0.0–1.0)

```

This structure lets downstream consumers work with typed objects rather than raw tuples or dictionaries.

## Implementation Example: Direct API Usage

For programmatic access to semantic search, import and call `semantic_code_search` directly:

```python
from codebase_rag.tools.semantic_search import semantic_code_search
from codebase_rag.services import QueryProtocol  # Neo4j/GraphQL ingestor

# Obtain a configured ingestor for your graph database

ingestor: QueryProtocol = ...  

results = semantic_code_search(
    ingestor,
    query="find all functions that parse JSON",
    top_k=5,
    project="my_project",  # Optional: restrict to specific project

)

for r in results:
    print(f"{r.qualified_name} (type={r.type}) – score {r.score:.3f}")

```

This low‑level approach gives you full control over result processing and integration with existing pipelines.

## Implementation Example: LLM Agent Integration

For agentic workflows, use `create_semantic_search_tool` to get a ready‑made, formatted interface:

```python
from codebase_rag.tools.semantic_search import create_semantic_search_tool

ingestor: QueryProtocol = ...  # Your graph ingestor

semantic_tool = create_semantic_search_tool(ingestor)

async def ask():
    response = await semantic_tool.run(
        "search for functions that open a file", 
        top_k=3
    )
    print(response)

# Running the coroutine produces formatted output:

# 1. my_module.file_io.open_file (type: function, score: 0.912)

# 2. utils.io.read_file (type: function, score: 0.887)

# 3. legacy.files.open_wrapper (type: function, score: 0.854)

```

The tool handles async execution, result enumeration, and human‑readable formatting automatically.

## Key Source Files

| File | Responsibility |
|------|---------------|
| [`codebase_rag/tools/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tools/semantic_search.py) | Orchestrates the full pipeline: embedding → lookup → formatting |
| [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py) | `embed_code` and batch variants for vector generation |
| [`codebase_rag/vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/vector_store.py) | `search_embeddings` abstracts vector store operations |
| [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py) | `SemanticSearchResult` data model |
| [`codebase_rag/utils/dependencies.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/dependencies.py) | `has_semantic_dependencies` guards optional features |

## Graceful Degradation

The implementation defensively checks for optional dependencies. When `has_semantic_dependencies` returns `False`—meaning embedding libraries or vector store clients aren't installed—`semantic_code_search` logs a warning and returns an empty list. This keeps the broader **code‑graph‑rag** system operational even in minimal installation environments.

## Summary

- **Query embedding** happens via `embed_code` in [`embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/embedder.py), supporting OpenAI or UnixCoder
- **Vector search** is handled by `search_embeddings` in [`vector_store.py`](https://github.com/vitali87/code-graph-rag/blob/main/vector_store.py) against Milvus Lite
- **Result assembly** occurs in `semantic_code_search` in [`tools/semantic_search.py`](https://github.com/vitali87/code-graph-rag/blob/main/tools/semantic_search.py)
- Use **direct API calls** for custom integrations or **`create_semantic_search_tool`** for LLM agents
- The system **gracefully degrades** when semantic dependencies are unavailable

## Frequently Asked Questions

### What embedding models does code‑graph‑rag support?

**code‑graph‑rag** supports two embedding backends: the **OpenAI embedding API** for cloud‑based operation and **UnixCoder**, a code‑specific model bundled with the repository for local execution. Both are accessible through the unified `embed_code` interface in [`codebase_rag/embedder.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/embedder.py).

### Can I use semantic search without installing all dependencies?

No—semantic search requires optional dependencies that are checked via `has_semantic_dependencies` in [`codebase_rag/utils/dependencies.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/dependencies.py). If these are missing, `semantic_code_search` returns an empty list with a logged warning, allowing the rest of the system to function normally.

### How are similarity scores calculated?

Scores come directly from the vector store's distance metric (typically **cosine similarity**), normalized to a 0.0–1.0 range. Higher scores indicate closer semantic alignment between your natural‑language query and the indexed code entity.

### What is the difference between `semantic_code_search` and `create_semantic_search_tool`?

**`semantic_code_search`** is the low‑level function returning structured `SemanticSearchResult` objects for programmatic processing. **`create_semantic_search_tool`** wraps this functionality into a formatted, async‑compatible interface suitable for LLM agents, handling string serialization and enumeration automatically.