# How Understand-Anything Captures Structural and Semantic Information in Codebases

> Discover how Understand Anything unifies code structure and semantics. Learn how tree-sitter parsing and LLM summaries create a powerful knowledge graph for Lum1104/Understand-Anything.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-01

---

**Understand-Anything captures both structural and semantic information by constructing a unified knowledge graph that combines tree-sitter parsing for code relationships with LLM-generated summaries for human-readable context.**

Every code comprehension tool faces the same challenge: understanding not just what files exist and how they connect, but also what they actually do. The Lum1104/Understand-Anything repository solves this by building a dual-layer knowledge graph that merges syntactic code analysis with natural language understanding. This approach enables precise querying, automated documentation, and contextual answers about any codebase.

## The Dual-Layer Knowledge Graph Architecture

The system treats structural and semantic information as orthogonal but complementary layers, storing both in a unified `GraphNode` structure.

### Structural Layer: Code Entities and Relationships

The **structural layer** captures the physical organization and dependencies within a repository. According to the source code in [`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts), the `GraphBuilder` class walks the repository and uses language-specific parsers via **tree-sitter** to produce a `StructuralAnalysis` object.

For every code entity discovered—whether files, functions, classes, imports, or service definitions—the builder creates a `GraphNode`. Relationships such as "contains", "imports", and "calls" are recorded as `GraphEdge` objects. Specific methods like `addImportEdge` and `addCallEdge` (found at lines 60-120 in graph-builder.ts) wire these connections together, creating a complete map of how code elements interact syntactically.

### Semantic Layer: LLM-Driven Analysis and Metadata

While the structural layer tells you *where* things are, the **semantic layer** explains *what* they do. This layer is populated by prompting an LLM to analyze each source file using `buildFileAnalysisPrompt` and the entire project via `buildProjectSummaryPrompt`.

As implemented in [`understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts) (lines 16-45), the system parses JSON responses through `parseFileAnalysisResponse` and `parseProjectSummaryResponse`. The extracted fields—including `summary`, `tags`, `complexity`, and `languageNotes`—are stored directly on the corresponding `GraphNode`, enriching the graph with human-readable context.

## Implementation: Building the Graph

The `GraphBuilder` API provides explicit methods to construct this dual-layer representation. You initialize a builder for a specific project, then incrementally add files with both their structural positions and semantic metadata:

```typescript
import { GraphBuilder } from "@understand-anything/core";

const builder = new GraphBuilder("my‑app", "a1b2c3d4");

// Add a plain file (no deep analysis)
builder.addFile("src/util.ts", {
  summary: "Utility helpers",
  tags: ["utility"],
  complexity: "simple",
});

// Add a file with structural analysis (functions & classes)
// `analysis` is the result of a language‑specific parser (tree‑sitter)
builder.addFileWithAnalysis("src/auth.ts", analysis, {
  fileSummary: "Authentication module",
  tags: ["auth"],
  complexity: "moderate",
  summaries: {
    login: "Handles user login",
    verify: "Verifies JWT tokens",
  },
});

// Record import and call relationships
builder.addImportEdge("src/auth.ts", "src/util.ts");
builder.addCallEdge("src/auth.ts", "login", "src/util.ts", "hashPassword");

// Final graph ready for LLM consumption
const graph = builder.build();

```

*Source:* `GraphBuilder` implementation – **graph-builder.ts** (lines 84-102).

## Enriching Nodes with Semantic Data

Bridging static code analysis with LLM inference requires careful prompt engineering and response parsing. The system constructs specialized prompts for file-level analysis and parses the structured output to populate node properties:

```typescript
import {
  buildFileAnalysisPrompt,
  parseFileAnalysisResponse,
} from "@understand-anything/core";

// Assume `fileContent` is the raw source string
const prompt = buildFileAnalysisPrompt(
  "src/auth.ts",
  fileContent,
  "A Node.js API that provides JWT authentication",
);

// Send `prompt` to the configured LLM (the plugin does this internally)
// `llmResponse` is the raw text returned by the model
const analysis = parseFileAnalysisResponse(llmResponse);
if (analysis) {
  // Use `analysis` to fill the node’s semantic fields
  builder.addFileWithAnalysis("src/auth.ts", structuralInfo, {
    fileSummary: analysis.fileSummary,
    tags: analysis.tags,
    complexity: analysis.complexity,
    summaries: analysis.functionSummaries,
    languageNotes: analysis.languageNotes,
  });
}

```

*Source:* LLM prompt construction and response parsing – **llm-analyzer.ts** (lines 16-45).

## Retrieval: Searching Across Both Dimensions

Once the graph is constructed, the `SearchEngine` indexes nodes using **Fuse.js**, applying specific weights to semantic fields. According to [`understand-anything-plugin/packages/core/src/search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/search.ts) (lines 14-20), the engine prioritizes `name`, `tags`, `summary`, and `languageNotes` when matching user queries.

After fuzzy matching identifies relevant nodes, the `buildChatContext` function (defined in [`understand-anything-plugin/src/context-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/src/context-builder.ts), lines 20-48) expands results by one hop via edges to include structurally related entities. Finally, `formatContextForPrompt` assembles these into formatted markdown that blends structural outlines with semantic annotations, ready for downstream LLM consumption.

```typescript
import { SearchEngine } from "@understand-anything/core";

const engine = new SearchEngine(graph.nodes);
const results = engine.search("jwt authentication");

// Expand one hop via edges (handled inside `buildChatContext`)
const chatContext = buildChatContext(graph, "How does the JWT flow work?");

// Render a markdown prompt for the LLM
const prompt = formatContextForPrompt(chatContext);

```

*Sources:* Search weighting – **search.ts** (lines 14-20); context assembly – **context-builder.ts** (lines 20-48).

## Summary

- **Structural information** is captured via tree-sitter parsing in `GraphBuilder`, which creates `GraphNode` and `GraphEdge` objects for files, functions, classes, and their relationships.
- **Semantic information** is generated through LLM analysis using prompts in [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts), populating nodes with summaries, tags, complexity ratings, and language-specific notes.
- **Unified retrieval** leverages Fuse.js to search across both dimensions, weighting semantic fields while maintaining structural context through graph traversal.
- **Final output** combines both layers into formatted markdown via `formatContextForPrompt`, enabling precise, context-aware answers about the codebase.

## Frequently Asked Questions

### What specific code relationships does Understand-Anything extract?

The system extracts import dependencies, function calls, class hierarchies, and containment relationships. In [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts), the `addImportEdge` method records file-to-file dependencies, while `addCallEdge` tracks specific function invocations between different source files, creating a comprehensive dependency graph.

### How does the LLM generate semantic summaries without hallucinating?

The system constrains LLM outputs through structured JSON schemas parsed by `parseFileAnalysisResponse` and `parseProjectSummaryResponse`. By providing the full file content as context and requesting specific fields (summary, tags, complexity, languageNotes), the prompts minimize hallucination and ensure responses align with actual code functionality.

### Can Understand-Anything work with any programming language?

Yes, the architecture is language-agnostic at the structural level due to **tree-sitter**, which provides parsers for most popular languages. The LLM semantic analysis is inherently language-agnostic, though `languageNotes` in the response can capture language-specific idioms and patterns detected in the source.

### How does the search engine balance structural vs semantic relevance?

The `SearchEngine` uses Fuse.js with explicit field weighting as shown in [`search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/search.ts) (lines 14-20), prioritizing semantic fields like `tags` and `summary` over raw file names. After semantic matching, the system expands results structurally via `buildChatContext` to include related nodes one edge away, ensuring both relevance and contextual completeness.