How Understand-Anything Captures Structural and Semantic Information in Codebases
Understand-Anything captures both structural and semantic information by constructing a unified knowledge graph that combines tree-sitter parsing for code relationships with LLM-generated summaries for human-readable context.
Every code comprehension tool faces the same challenge: understanding not just what files exist and how they connect, but also what they actually do. The Lum1104/Understand-Anything repository solves this by building a dual-layer knowledge graph that merges syntactic code analysis with natural language understanding. This approach enables precise querying, automated documentation, and contextual answers about any codebase.
The Dual-Layer Knowledge Graph Architecture
The system treats structural and semantic information as orthogonal but complementary layers, storing both in a unified GraphNode structure.
Structural Layer: Code Entities and Relationships
The structural layer captures the physical organization and dependencies within a repository. According to the source code in understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts, the GraphBuilder class walks the repository and uses language-specific parsers via tree-sitter to produce a StructuralAnalysis object.
For every code entity discovered—whether files, functions, classes, imports, or service definitions—the builder creates a GraphNode. Relationships such as "contains", "imports", and "calls" are recorded as GraphEdge objects. Specific methods like addImportEdge and addCallEdge (found at lines 60-120 in graph-builder.ts) wire these connections together, creating a complete map of how code elements interact syntactically.
Semantic Layer: LLM-Driven Analysis and Metadata
While the structural layer tells you where things are, the semantic layer explains what they do. This layer is populated by prompting an LLM to analyze each source file using buildFileAnalysisPrompt and the entire project via buildProjectSummaryPrompt.
As implemented in understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts (lines 16-45), the system parses JSON responses through parseFileAnalysisResponse and parseProjectSummaryResponse. The extracted fields—including summary, tags, complexity, and languageNotes—are stored directly on the corresponding GraphNode, enriching the graph with human-readable context.
Implementation: Building the Graph
The GraphBuilder API provides explicit methods to construct this dual-layer representation. You initialize a builder for a specific project, then incrementally add files with both their structural positions and semantic metadata:
import { GraphBuilder } from "@understand-anything/core";
const builder = new GraphBuilder("my‑app", "a1b2c3d4");
// Add a plain file (no deep analysis)
builder.addFile("src/util.ts", {
summary: "Utility helpers",
tags: ["utility"],
complexity: "simple",
});
// Add a file with structural analysis (functions & classes)
// `analysis` is the result of a language‑specific parser (tree‑sitter)
builder.addFileWithAnalysis("src/auth.ts", analysis, {
fileSummary: "Authentication module",
tags: ["auth"],
complexity: "moderate",
summaries: {
login: "Handles user login",
verify: "Verifies JWT tokens",
},
});
// Record import and call relationships
builder.addImportEdge("src/auth.ts", "src/util.ts");
builder.addCallEdge("src/auth.ts", "login", "src/util.ts", "hashPassword");
// Final graph ready for LLM consumption
const graph = builder.build();
Source: GraphBuilder implementation – graph-builder.ts (lines 84-102).
Enriching Nodes with Semantic Data
Bridging static code analysis with LLM inference requires careful prompt engineering and response parsing. The system constructs specialized prompts for file-level analysis and parses the structured output to populate node properties:
import {
buildFileAnalysisPrompt,
parseFileAnalysisResponse,
} from "@understand-anything/core";
// Assume `fileContent` is the raw source string
const prompt = buildFileAnalysisPrompt(
"src/auth.ts",
fileContent,
"A Node.js API that provides JWT authentication",
);
// Send `prompt` to the configured LLM (the plugin does this internally)
// `llmResponse` is the raw text returned by the model
const analysis = parseFileAnalysisResponse(llmResponse);
if (analysis) {
// Use `analysis` to fill the node’s semantic fields
builder.addFileWithAnalysis("src/auth.ts", structuralInfo, {
fileSummary: analysis.fileSummary,
tags: analysis.tags,
complexity: analysis.complexity,
summaries: analysis.functionSummaries,
languageNotes: analysis.languageNotes,
});
}
Source: LLM prompt construction and response parsing – llm-analyzer.ts (lines 16-45).
Retrieval: Searching Across Both Dimensions
Once the graph is constructed, the SearchEngine indexes nodes using Fuse.js, applying specific weights to semantic fields. According to understand-anything-plugin/packages/core/src/search.ts (lines 14-20), the engine prioritizes name, tags, summary, and languageNotes when matching user queries.
After fuzzy matching identifies relevant nodes, the buildChatContext function (defined in understand-anything-plugin/src/context-builder.ts, lines 20-48) expands results by one hop via edges to include structurally related entities. Finally, formatContextForPrompt assembles these into formatted markdown that blends structural outlines with semantic annotations, ready for downstream LLM consumption.
import { SearchEngine } from "@understand-anything/core";
const engine = new SearchEngine(graph.nodes);
const results = engine.search("jwt authentication");
// Expand one hop via edges (handled inside `buildChatContext`)
const chatContext = buildChatContext(graph, "How does the JWT flow work?");
// Render a markdown prompt for the LLM
const prompt = formatContextForPrompt(chatContext);
Sources: Search weighting – search.ts (lines 14-20); context assembly – context-builder.ts (lines 20-48).
Summary
- Structural information is captured via tree-sitter parsing in
GraphBuilder, which createsGraphNodeandGraphEdgeobjects for files, functions, classes, and their relationships. - Semantic information is generated through LLM analysis using prompts in
llm-analyzer.ts, populating nodes with summaries, tags, complexity ratings, and language-specific notes. - Unified retrieval leverages Fuse.js to search across both dimensions, weighting semantic fields while maintaining structural context through graph traversal.
- Final output combines both layers into formatted markdown via
formatContextForPrompt, enabling precise, context-aware answers about the codebase.
Frequently Asked Questions
What specific code relationships does Understand-Anything extract?
The system extracts import dependencies, function calls, class hierarchies, and containment relationships. In graph-builder.ts, the addImportEdge method records file-to-file dependencies, while addCallEdge tracks specific function invocations between different source files, creating a comprehensive dependency graph.
How does the LLM generate semantic summaries without hallucinating?
The system constrains LLM outputs through structured JSON schemas parsed by parseFileAnalysisResponse and parseProjectSummaryResponse. By providing the full file content as context and requesting specific fields (summary, tags, complexity, languageNotes), the prompts minimize hallucination and ensure responses align with actual code functionality.
Can Understand-Anything work with any programming language?
Yes, the architecture is language-agnostic at the structural level due to tree-sitter, which provides parsers for most popular languages. The LLM semantic analysis is inherently language-agnostic, though languageNotes in the response can capture language-specific idioms and patterns detected in the source.
How does the search engine balance structural vs semantic relevance?
The SearchEngine uses Fuse.js with explicit field weighting as shown in search.ts (lines 14-20), prioritizing semantic fields like tags and summary over raw file names. After semantic matching, the system expands results structurally via buildChatContext to include related nodes one edge away, ensuring both relevance and contextual completeness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →