How Fuzzy and Semantic Search Find Code by Meaning in Understand Anything

Understand Anything combines Fuse.js for weighted fuzzy matching against node metadata with cosine similarity on vector embeddings for semantic retrieval, returning unified scores that let users toggle between keyword and meaning-based code search.

Understand Anything analyzes codebases by constructing a knowledge graph where every code artifact—functions, classes, files, and more—is represented as a GraphNode defined in packages/core/src/types.ts. To enable developers to locate code by intent rather than exact identifiers alone, the system implements two complementary search strategies that operate on this unified graph: a fuzzy search engine for approximate keyword matching and a semantic search engine for conceptual vector similarity.

Fuzzy Search with Fuse.js

The fuzzy search capability is implemented in the SearchEngine class located at packages/core/src/search.ts. This engine leverages the Fuse.js library to perform approximate string matching across node metadata, making it resilient to typos and partial word matches while biasing results toward the most informative fields.

Indexing Strategy and Weighted Fields

At construction, the engine receives all graph nodes and creates a Fuse.js index via new Fuse(nodes, FUSE_OPTIONS). The FUSE_OPTIONS object assigns specific weights to fields based on their semantic importance: name receives 0.4, tags 0.3, summary 0.2, and languageNotes 0.1. This weighting ensures that matches against identifiers and tagged concepts rank higher than matches in descriptive text. The configuration also sets threshold: 0.4 to discard low-quality matches, enables ignoreLocation: true to find matches anywhere in the text, and activates useExtendedSearch: true to support token-based queries.

Query Processing and Tokenization

When processing a query, the engine transforms the raw input to support OR-style matching. The implementation trims the query, splits it on whitespace, and rejoins tokens with the | operator. For example, the query "auth contrl" becomes "auth | contrl" (see line 46 in search.ts). Fuse.js returns results with a score where 0 indicates a perfect match and 1 indicates the worst match. The engine maps these into an array of { nodeId, score } objects, identical to the semantic engine’s output format. Optional SearchOptions.types parameters can filter results to specific node categories such as function or class.

// Fuzzy search – find nodes that mention "authentication"
import { SearchEngine } from "@understand-anything/core";
import type { GraphNode } from "@understand-anything/core";

const nodes: GraphNode[] = …;               // the graph’s nodes
const fuzzy = new SearchEngine(nodes);
const fuzzyResults = fuzzy.search("authentication");

// fuzzyResults → [{ nodeId: "function:auth.ts:login", score: 0.12 }, …]

Semantic Search with Vector Embeddings

For meaning-based retrieval, Understand Anything provides the SemanticSearchEngine class in packages/core/src/embedding-search.ts. This engine assumes each GraphNode has been enriched with a pre-computed vector embedding—typically generated via an external LLM API—and stores these in a Map<string, number[]> keyed by node ID.

Cosine Similarity Calculation

The core similarity metric is implemented in the cosineSimilarity function (lines 14-30), which computes the cosine of the angle between two vectors, returning a value in the range [0, 1] where 1 represents identical direction and 0 represents orthogonality. During a search, the engine iterates through all nodes with embeddings, calculates similarity = cosineSimilarity(queryEmbedding, nodeEmbedding), and filters candidates against a configurable threshold (default 0).

Score Normalization and Unification

To maintain interface consistency with the fuzzy engine, semantic scores undergo inversion: the engine stores { nodeId, score: 1 - similarity } so that lower scores consistently indicate better matches across both search modes. Results are sorted by ascending score and limited to the top N entries (default 10), with optional types filtering to restrict results to specific node categories.

// Semantic search – find code handling authentication even if the word isn’t present
import { SemanticSearchEngine } from "@understand-anything/core";

// Assume each node already has an embedding (generated by an external LLM API)
const embeddings: Record<string, number[]> = {
  "function:auth.ts:login": [0.12, 0.03, …],
  "class:User.ts": [0.04, 0.19, …],
  // …
};

const semantic = new SemanticSearchEngine(nodes, embeddings);

// `queryEmbedding` comes from the same embedding model used for nodes
const queryEmbedding: number[] = [0.10, 0.05, …];
const semanticResults = semantic.search(queryEmbedding, { limit: 5 });

// semanticResults → sorted by 1‑similarity (lower score = closer match)

Dashboard Integration and Mode Switching

The unified search interface is managed in the dashboard state at packages/dashboard/src/store.ts. The searchMode state determines which engine instance handles the query: when semantic mode is active and embeddings are present, the system uses SemanticSearchEngine; otherwise it falls back to the fuzzy SearchEngine (see lines 528-532). Both engines implement an identical search method signature, returning { nodeId, score } arrays that allow the UI to toggle between "Fuzzy" and "Semantic" modes without changing downstream display logic.

Higher-level features consume these search results through utilities like buildChatContext in understand-anything-plugin/src/context-builder.ts. This helper internally instantiates a SearchEngine, retrieves top results, and expands the result set by one graph hop to provide the LLM with a concise, meaning-oriented view of the relevant codebase.

// Using the built‑in helper to build a chat context (fuzzy or semantic)
import { buildChatContext } from "@understand-anything/plugin";
import { SearchEngine } from "@understand-anything/core";

const graph = …;                     // full KnowledgeGraph
const query = "how does the app validate a token?";

const context = buildChatContext(graph, query);
// Internally `buildChatContext` creates a SearchEngine and expands the
// result set by one hop, giving the LLM a concise, meaning‑oriented view.

Summary

  • Understand Anything uses a dual-engine architecture combining Fuse.js fuzzy matching and embedding-based cosine similarity to search its knowledge graph of GraphNode objects.
  • The fuzzy engine in packages/core/src/search.ts applies weighted field matching (name 0.4, tags 0.3, summary 0.2) with a 0.4 threshold and OR-tokenized queries to handle typos.
  • The semantic engine in packages/core/src/embedding-search.ts computes cosine similarity between query and node embeddings, inverting scores to { nodeId, 1-similarity } for unified ranking.
  • Both engines return identical result shapes, enabling seamless toggling between keyword and meaning-based search modes in the dashboard without UI code changes.

Frequently Asked Questions

What is the difference between fuzzy and semantic search in Understand Anything?

Fuzzy search matches keywords against node metadata using approximate string matching with weighted fields, making it ideal for finding known identifiers despite typos or partial matches. Semantic search compares vector embeddings via cosine similarity to find conceptually related code even when the query keywords differ entirely from the source text.

How does the fuzzy search handle typos in queries?

The Fuse.js implementation in SearchEngine automatically handles typos through approximate matching algorithms, further enhanced by the ignoreLocation option and the OR-tokenization strategy that splits queries like "auth contrl" into "auth | contrl" to match partial terms across different fields.

Can I customize the search weights or thresholds?

The FUSE_OPTIONS in packages/core/src/search.ts hardcode field weights (name 0.4, tags 0.3, etc.) and a threshold of 0.4, while the semantic engine accepts a configurable threshold parameter (default 0) in its search method. Advanced customization would require modifying these configuration objects in the source code.

Where do the embeddings for semantic search come from?

The SemanticSearchEngine expects pre-computed embeddings stored in a Map<string, number[]>, typically generated by an external LLM embedding API before the search operation. The engine itself focuses purely on similarity calculation and does not generate embeddings internally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →