# How Understand Anything Architecture Works: A Deep Dive into the Knowledge Graph System

> Discover how Understand Anything's monorepo architecture works. Learn about its JSON knowledge graph, incremental analysis, and LLM enrichment for deep codebase insights.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-05

---

**Understand Anything uses a monorepo architecture that separates analysis (core engine and Claude Code skill) from visualization (React dashboard) through a JSON knowledge graph file, enabling incremental codebase analysis via tree-sitter parsing and LLM enrichment.**

Understand Anything is an open-source tool designed to help developers comprehend complex codebases through automated analysis and interactive visualization. The system employs a modular architecture that processes source code into a structured knowledge graph, enabling both command-line interaction and rich UI exploration. This article examines the technical architecture based on the implementation in the `Lum1104/Understand-Anything` repository.

## Monorepo Structure and Component Separation

The Understand Anything architecture is organized as a **pnpm workspace monorepo** that isolates three distinct runtimes. This separation ensures the heavy analysis engine remains framework-agnostic while the dashboard stays lightweight and browser-compatible.

The three primary components communicate exclusively through a **JSON knowledge graph file** ([`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json)):

- **Core Engine** (`packages/core`): A shared library handling static analysis, LLM-driven reasoning, and graph construction. This package is imported by both the skill and dashboard layers without pulling in browser-only modules.
- **Claude Code Skill** (`src/`): A command-line interface providing slash commands like `/understand` and `/understand-chat`. It drives the analysis pipeline and writes results to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json).
- **Dashboard** (`packages/dashboard`): A React and TypeScript visualization layer that renders the knowledge graph using React Flow, provides search capabilities, and hosts interactive learning tours.

The design specification in [`docs/superpowers/specs/2026-03-14-understand-anything-design.md`](https://github.com/Lum1104/Understand-Anything/blob/main/docs/superpowers/specs/2026-03-14-understand-anything-design.md) documents this architecture, noting that the JSON interchange isolates the UI from the heavy analysis stack. This allows the dashboard to run completely offline once the graph file exists.

## Knowledge Graph Schema and Communication Protocol

The knowledge graph schema serves as the **contract** between the analysis engine and visualization layer. Defined in [`packages/core/src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/schema.ts), the TypeScript schema ensures type safety across the monorepo while enabling language-agnostic data exchange.

Key types include:

- `KnowledgeGraph`: The top-level container storing version metadata, nodes, edges, layers, and tours.
- `GraphNode`: Represents files, functions, classes, modules, or abstract concepts.
- `GraphEdge`: Defines relationships such as imports, calls, or dependencies.
- `Layer`: Logical groupings like "API Layer" or "Data Layer" for architectural visualization.
- `TourStep`: Guided walkthrough steps for the dashboard's "Learn" panel.

This schema enables the dashboard to function as a **stateless consumer** of the analysis data. By reading the static JSON file, the dashboard avoids direct dependencies on the analysis engine, parser libraries, or LLM clients.

## Core Analysis Pipeline

The analysis flow orchestrates multiple specialized agents to transform raw source code into a structured graph. This pipeline is implemented in the core engine and invoked by the skill layer through [`src/understand-chat.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/understand-chat.ts) and [`src/onboard-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/onboard-builder.ts).

### File Discovery and Structural Parsing

The process begins with a **project-scanner** agent that walks the repository and identifies relevant files. These files are passed to the tree-sitter plugin system located in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts).

The tree-sitter plugin extracts structural elements including function definitions, class hierarchies, and import statements. This structural analysis operates independently of LLM processing, ensuring fast, deterministic parsing of code syntax for supported languages.

### LLM Enrichment and Summarization

After structural parsing, the **llm-analyzer** (located in [`packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/llm-analyzer.ts)) processes file snippets through the Claude API. This enrichment phase generates plain-English summaries of complex functions, semantic tags for categorization, and language-specific implementation notes.

The LLM layer adds **semantic understanding** that static analysis alone cannot capture, such as the architectural intent behind a module or the business logic purpose of a specific function.

### Graph Construction and Persistence

The **graph-builder** ([`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts)) merges structural data from tree-sitter with semantic data from the LLM analyzer. It constructs the final `KnowledgeGraph` object with validated nodes and edges.

Persistence is handled by the fingerprinting system in [`core/src/persistence/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/core/src/persistence/fingerprint.ts), which writes the graph to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) and manages staleness metadata. This file serves as the **single source of truth** for the entire system.

## Dashboard Visualization Layer

The dashboard (`packages/dashboard`) is a React application that renders the knowledge graph using specialized visualization libraries. It consumes the static JSON file produced by the core engine.

Key implementation details include:

- **React Flow** renders the interactive node graph, allowing users to explore dependencies visually.
- **Monaco Editor** provides read-only code views with syntax highlighting.
- **Zustand** manages global state for search queries, selected nodes, and UI panels.

The entry point [`packages/dashboard/src/App.tsx`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/dashboard/src/App.tsx) sets up the panel layout, while [`packages/dashboard/src/hooks/useKeyboardShortcuts.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/dashboard/src/hooks/useKeyboardShortcuts.ts) handles global hotkeys like search focus. The dashboard can run in **stand-alone mode** (requiring a Claude API key for chat features) or **offline mode** (graph search and exploration work without network access).

## Incremental Updates and Staleness Detection

To maintain performance on large codebases, Understand Anything implements **incremental analysis** based on git state. The staleness detection system is implemented in `core/src/persistence/` and operates as follows:

1. Reads [`meta.json`](https://github.com/Lum1104/Understand-Anything/blob/main/meta.json) to retrieve the last analyzed commit hash.
2. Executes `git diff <last-hash>..HEAD --name-only` to identify changed files.
3. Re-analyzes only modified files, merging new nodes and edges into the existing graph.
4. Updates [`meta.json`](https://github.com/Lum1104/Understand-Anything/blob/main/meta.json) with the new commit hash and timestamps.

This approach ensures that running `/understand` after a small code change completes in seconds rather than minutes, as only the delta requires reprocessing.

## Plugin System for Language Support

The core engine exposes an **`AnalyzerPlugin`** interface that enables extensible language support. Built-in plugins reside in `plugins/tree-sitter` and provide parsers for TypeScript, Python, Go, Java, Rust, and C/C++.

The interface, documented in [`docs/superpowers/specs/2026-03-14-understand-anything-design.md`](https://github.com/Lum1104/Understand-Anything/blob/main/docs/superpowers/specs/2026-03-14-understand-anything-design.md), requires:

```typescript
interface AnalyzerPlugin {
  name: string;
  languages: string[];
  analyzeFile(filePath: string, content: string): StructuralAnalysis;
  resolveImports(filePath: string, content: string): ImportResolution[];
  extractCallGraph?(filePath: string, content: string): CallGraphEntry[];
}

```

Community plugins can implement this interface and be dropped into the `plugins/` directory without modifying core engine code. The tree-sitter plugin ([`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts)) serves as the reference implementation, using Web-Tree-Sitter for language-agnostic parsing.

## Summary

The Understand Anything architecture achieves loose coupling between analysis and visualization through a JSON knowledge graph intermediary. Key architectural decisions include:

- **Monorepo organization** using pnpm workspaces to share the core engine between CLI and UI components while maintaining separation of concerns.
- **JSON-based communication** enabling offline dashboard operation and language-agnostic data exchange via the schema defined in [`packages/core/src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/schema.ts).
- **Incremental analysis** via git-diff staleness detection in `core/src/persistence/`, minimizing reprocessing on large codebases.
- **Plugin architecture** supporting extensible language parsing through the standardized `AnalyzerPlugin` TypeScript interface.
- **LLM-augmented analysis** combining deterministic tree-sitter parsing with semantic enrichment from Claude.

## Frequently Asked Questions

### How does Understand Anything handle large codebases efficiently?

The system uses **incremental analysis** based on git diff detection stored in [`core/src/persistence/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/core/src/persistence/fingerprint.ts). When you run the `/understand` command, it compares the current commit hash against the last analyzed state stored in [`meta.json`](https://github.com/Lum1104/Understand-Anything/blob/main/meta.json). Only files modified since the last analysis are reprocessed and merged into the existing knowledge graph, reducing analysis time from minutes to seconds on subsequent runs.

### What file format does Understand Anything use to store analysis results?

Understand Anything stores all analysis results in a **JSON knowledge graph file** located at [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). This file follows the strict TypeScript schema defined in [`packages/core/src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/schema.ts) and contains nodes (files, functions, classes), edges (relationships), layers (architectural groupings), and tour steps. This format allows the dashboard to function independently of the analysis engine.

### Can the Understand Anything dashboard run without an internet connection?

Yes. The dashboard can operate in **offline mode** for graph search, node exploration, and code viewing once the [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) file exists. However, the chat features and LLM-powered explainers require a Claude API key and network access. The React-based UI uses React Flow for visualization and Monaco Editor for code display without external dependencies.

### How can I add support for a new programming language to Understand Anything?

You can implement the **`AnalyzerPlugin`** interface and place your plugin in the `plugins/` directory. Your implementation must provide `analyzeFile` for structural extraction, `resolveImports` for dependency mapping, and optionally `extractCallGraph` for call relationship tracking. The core engine dynamically loads these plugins, allowing language support extensions without modifying the core analysis logic in `packages/core/src/analyzer/`.