How Understand-Anything Implements Semantic Search Using Embeddings to Find Relevant Code
Understand-Anything implements semantic search by generating dense vector embeddings for each code node via LLM agents, storing them in a SemanticSearchEngine, and ranking results using cosine similarity between the query embedding and stored node embeddings.
The Understand-Anything project adds a semantic layer to code exploration by mapping natural language queries to code elements through vector embeddings. Unlike traditional keyword search, this approach understands the meaning behind both the query and the code structure. This article examines how the system implements semantic search using embeddings to surface relevant code based on conceptual similarity rather than literal string matches.
How Semantic Search Works in Understand-Anything
Step 1: Embedding Generation via LLM Agents
When the analysis pipeline runs, LLM-based agents (specifically the file-analyzer and architecture-analyzer) generate a dense vector embedding for each graph node. These embeddings encode semantic meaning—such as "authentication flow" or "database helper"—rather than just the node's literal name. The vectors are stored in the graph's persisted JSON under the embeddings field, as referenced in the EmbeddingSearchEngine constructor.
Step 2: Vector Indexing in SemanticSearchEngine
The SemanticSearchEngine class, defined in understand-anything-plugin/packages/core/src/embedding-search.ts, receives the full list of GraphNode objects alongside a mapping of node IDs to their embedding vectors. It maintains these vectors in a Map<string, number[]> structure, enabling O(1) lookup when performing similarity calculations.
Step 3: Cosine Similarity Ranking
When a user submits a natural language query (e.g., "how is auth handled?"), the system converts this into an embedding using the same LLM. The engine then iterates over every indexed node, retrieves its stored embedding, and computes the cosine similarity using the cosineSimilarity function:
const similarity = cosineSimilarity(queryEmbedding, nodeEmbedding);
Nodes exceeding a configurable threshold are collected, converted to a distance score (1 - similarity where lower is better), sorted, and returned as SearchResult[] objects. Optional types filters allow restricting results to specific node types like functions or classes.
Core Components and File Structure
Understanding the architecture requires examining specific source files:
packages/core/src/embedding-search.ts: Contains theSemanticSearchEngineclass andcosineSimilarityfunction. This is the core implementation where vector comparison and ranking occur.packages/core/src/types.ts: Defines theGraphNodeinterface, which pairs node IDs with type metadata and semantic information.packages/dashboard/src/store.ts: Integration point where the UI switches between fuzzy text search and semantic search based on embedding availability.packages/core/src/__tests__/embedding-search.test.ts: Unit tests verifying distance calculations and search behavior.
Implementing Semantic Search: Code Examples
Initializing the Search Engine
To create a search instance, provide the graph nodes and pre-computed embeddings:
import { SemanticSearchEngine } from "@understand-anything/core/embedding-search";
import type { GraphNode } from "@understand-anything/core/types";
const graphNodes: GraphNode[] = // ... loaded from knowledge graph
const initialEmbeddings = {
"node-1": [0.12, 0.34, 0.56],
"node-2": [0.88, 0.01, 0.43],
};
const engine = new SemanticSearchEngine(graphNodes, initialEmbeddings);
// Add incremental embeddings as analysis progresses
engine.addEmbedding("node-42", [0.44, 0.22, 0.11]);
Executing a Semantic Query
Convert user queries to embeddings and search:
// Obtain embedding from LLM for the natural language query
const queryEmbedding = await llm.embed("Which parts handle authentication?");
// Search with filters and thresholds
const results = engine.search(queryEmbedding, {
limit: 5,
threshold: 0.6,
types: ["function", "class"]
});
results.forEach(({ nodeId, score }) => {
console.log(`Node ${nodeId} – distance ${score.toFixed(3)}`);
});
Dashboard Integration
The dashboard checks for embedding availability before enabling semantic mode:
if (searchEngine.hasEmbeddings()) {
const hits = searchEngine.search(userQueryEmbedding, { limit: 10 });
displaySemanticResults(hits);
}
Summary
- Embedding Generation: LLM agents create dense vectors encoding semantic meaning for each code node during the analysis pipeline.
- Vector Storage: The
SemanticSearchEnginemaintains embeddings in aMap<string, number[]>for O(1) access, initialized from the graph's persisted JSON. - Similarity Calculation: Cosine similarity in
cosineSimilaritycompares query embeddings against node embeddings to find conceptually related code. - Configurable Filtering: The search supports thresholds and type filters (e.g., functions only) to refine results.
- UI Integration: The dashboard in
store.tsautomatically switches to semantic search mode when embeddings are available.
Frequently Asked Questions
How does Understand-Anything generate embeddings for code nodes?
The system uses LLM-based agents (file-analyzer and architecture-analyzer) to generate dense vector embeddings during the pipeline execution. These embeddings capture semantic meaning like "authentication flow" rather than just literal text, and are stored in the graph's JSON under the embeddings field.
What similarity metric does the SemanticSearchEngine use?
The engine uses cosine similarity to measure the angle between two embedding vectors. The implementation in src/embedding-search.ts calculates this value, then converts it to a distance score (1 - similarity) for ranking, where lower scores indicate higher relevance.
Can I filter semantic search results by specific code types?
Yes. The search() method accepts an optional types parameter (e.g., ["function", "class"]) that restricts results to specific node types. This filtering occurs after similarity calculation but before returning the final ranked list.
Where does the dashboard decide between fuzzy and semantic search?
The decision logic resides in packages/dashboard/src/store.ts. The UI checks searchEngine.hasEmbeddings() to determine if semantic search is available, switching to "semantic" mode when embeddings exist and falling back to traditional text search otherwise.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →