How Fuzzy and Semantic Search Work in the Egonex-AI Knowledge Graph
The Understand-Anything plugin implements dual search modes—fuzzy matching via Fuse.js for typo-tolerant text search, and semantic matching via vector embeddings for meaning-based discovery—both managed through a unified searchMode state in the dashboard store.
The Egonex-AI/Understand-Anything repository builds an intelligent knowledge graph from codebase analysis, exposing two complementary search strategies that help developers navigate complex software structures. While fuzzy search locates nodes by approximate string matches against names and tags, semantic search retrieves conceptually related nodes using vector similarity, enabling both precise lookups and exploratory discovery.
Fuzzy Search Implementation
The fuzzy search engine resides in understand-anything-plugin/packages/core/src/search.ts, where the SearchEngine class wraps the Fuse.js library to provide fast, tolerant text matching across the knowledge graph.
How Fuse.js Matches Nodes
The engine indexes four searchable fields on every graph node: name, tags, summary, and languageNotes. It instantiates Fuse with FUSE_OPTIONS that assign specific weights to these fields, ensuring that matches in critical identifiers rank higher than those in supplementary notes.
Fuse.js scores matches from 0 (perfect match) to 1 (worst match), allowing the engine to return results even when queries contain typos or partial terms. When you invoke search(), the method first trims the input, then constructs an extended OR-search query where spaces become pipe operators—transforming "auth contrl" into "auth | contrl"—before executing fuse.search().
Filtering and Results
You can optionally restrict results by node type (e.g., function, class) and limit the output count. The method returns an array of SearchResult objects with a uniform interface:
interface SearchResult {
nodeId: string;
score: number; // Fuse.js distance (lower is better)
}
Semantic Search Implementation
For meaning-based retrieval, the repository provides SemanticSearchEngine in understand-anything-plugin/packages/core/src/embedding-search.ts. This engine operates on pre-computed vector embeddings that capture the semantic essence of each node's documentation and code context.
Vector Embeddings and Cosine Similarity
The engine stores embeddings in a Map<string, number[]> mapping node IDs to float arrays. When you submit a query, the system first converts the query text into an embedding vector via an external language model (the plugin manages only the vector storage and comparison).
The cosineSimilarity function calculates the angular distance between the query vector and each stored node vector, ranking nodes by similarity. Unlike fuzzy scoring, semantic scores range closer to 1 for highly similar vectors and approach 0 for dissimilar content. The top-k results are returned in the same { nodeId, score } format, though here a higher score indicates stronger semantic alignment.
Switching Between Search Modes
Both engines feed into a unified dashboard UI controlled by understand-anything-plugin/packages/dashboard/src/store.ts. The global Zustand store exposes a searchMode state that toggles between "fuzzy" and "semantic", along with a setSearchMode action to update it:
// store.ts excerpt
export const useStore = create<State>((set) => ({
searchMode: "fuzzy" as "fuzzy" | "semantic",
setSearchMode: (mode) => set({ searchMode: mode }),
// ...
}));
When the user enters a query, the dashboard checks store.searchMode and delegates to the appropriate engine. Because both SearchEngine and SemanticSearchEngine implement the same output interface, downstream components—including the graph viewer and sidebar—render results uniformly without mode-specific logic.
Practical Code Examples
Performing a Fuzzy Search
import { SearchEngine } from "@understand-anything/core/search";
import type { GraphNode } from "@understand-anything/core/types";
// Initialize with all nodes from the knowledge graph
const fuzzyEngine = new SearchEngine(graphNodes);
// Search for authentication controllers with typo tolerance
const results = fuzzyEngine.search("auth contrl", {
types: ["function", "class"],
limit: 10
});
// Results: [{ nodeId: "node-123", score: 0.12 }, ...]
Performing a Semantic Search
import { SemanticSearchEngine } from "@understand-anything/core/embedding-search";
// embeddings: Map<string, number[]> pre-computed by an LLM
const semanticEngine = new SemanticSearchEngine(graphNodes, embeddings);
// Discover nodes conceptually related to authentication
const results = semanticEngine.search(
"How does authentication work in this project?",
{ limit: 5 }
);
// Results: [{ nodeId: "node-456", score: 0.89 }, ...]
Toggling Search Mode in the UI
import { useStore } from "./store";
function toggleSearchMode() {
const store = useStore.getState();
const newMode = store.searchMode === "fuzzy" ? "semantic" : "fuzzy";
store.setSearchMode(newMode);
}
Summary
SearchEngineinsearch.tsimplements fuzzy matching using Fuse.js against node names, tags, summaries, and language notes, supporting weighted field prioritization and extended OR-queries.SemanticSearchEngineinembedding-search.tsenables meaning-based discovery through cosine similarity of vector embeddings, returning semantically related nodes regardless of keyword overlap.- The dashboard store in
store.tsmanages the activesearchModestate, allowing seamless switching between engines while maintaining a consistentSearchResultinterface for UI components. - Both engines support filtering by node type and result limiting, but differ in scoring semantics: fuzzy uses distance (lower is better) while semantic uses similarity (higher is better).
Frequently Asked Questions
How does the fuzzy search handle typos and partial matches?
The fuzzy search transforms user queries into extended OR-patterns (e.g., "auth contrl" becomes "auth | contrl") and leverages Fuse.js's bitap algorithm to calculate match distances across weighted fields. This allows it to return relevant nodes even when the query contains spelling errors or incomplete words, scoring results from 0 (exact match) to 1 (no match).
What powers the semantic search embeddings?
The semantic search engine stores pre-computed vector embeddings in a Map<string, number[]>, but does not generate them internally. An external language model produces these vectors by analyzing node content and documentation. The plugin then uses cosine similarity to compare the query vector against stored node vectors, ranking results by semantic proximity.
Can I filter search results by specific node types?
Yes. Both SearchEngine and SemanticSearchEngine accept an options object that includes a types array. When provided, the fuzzy engine filters the Fuse results post-search, while the semantic engine filters the candidate set before calculating similarities, allowing you to restrict results to functions, classes, or other graph node categories.
Where is the search mode preference stored?
The active search mode persists in the Zustand global store defined in understand-anything-plugin/packages/dashboard/src/store.ts. The useStore hook exposes searchMode and setSearchMode, which the UI uses to toggle between fuzzy and semantic strategies while maintaining the same result rendering logic downstream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →