# What Kind of Data Does Understand Anything Process? Sources, Node Types, and Code Examples

> Understand Anything processes structured text like source code and config files to build a unified Knowledge Graph. Discover its data sources, node types, and code examples.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-05

---

**Understand Anything ingests any structured text source—including source code, configuration files, documentation, and business-domain descriptors—to build a unified Knowledge Graph that mixes deterministic syntax extraction with LLM-generated semantic meaning.**

The open-source repository `Lum1104/Understand-Anything` is designed to transform any project artifact that can be represented as text into a rich, queryable graph. What kind of data Understand Anything processes spans from raw source files parsed by Tree-sitter to markdown wikis analyzed by LLM agents. All extracted nodes and edges are merged into a single Knowledge Graph at [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json), which powers search, diff-impact analysis, and guided tours.

## Source Code Files Parsed by Tree-sitter

At the core of the pipeline, the **Tree-sitter plugin** scans every file in a project and parses it into a concrete syntax tree. In [`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts) (lines 21-30), the plugin extracts **functions, classes, imports/exports, and call-graph edges** regardless of language. This structural layer is fully deterministic: the same code always yields the same nodes and relationships.

## Configuration and Build Artifacts

Beyond source code, Understand Anything processes infrastructure and build files through dedicated parsers. The index at [`understand-anything-plugin/packages/core/src/plugins/parsers/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/index.ts) (lines 4-17) registers handlers for **JSON, YAML, TOML, Dockerfile, SQL, GraphQL, protobuf, Terraform, Makefile, Shell, and Env** files. These parsers turn configuration data into structured nodes, allowing the graph to capture **configuration-driven dependencies** alongside code relationships.

## Documentation and Knowledge Bases

The pipeline treats documentation as first-class data. According to [`understand-anything-plugin/README.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/README.md) (lines 65-68), markdown-based wikis—such as “Karpathy-pattern” LLM wikis—are fed to the **article-analyzer** agent. This agent extracts **entities, claims, and implicit relationships** to build a knowledge graph of ideas that sits alongside the technical graph.

## Semantic Layers: Business Domains and Language Concepts

After structural extraction, specialized analyzers annotate the graph with higher-level meaning.

### Business-Domain Tagging

The **domain-analyzer** agent—centralized in [`understand-anything-plugin/packages/core/src/analyzer/domain-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/domain-analyzer.ts) and orchestrated within [`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts) (lines 89-97)—tags code with **domains, flows, and process steps**. This turns technical artifacts into business-oriented concepts that are stored as `domain` nodes in the final graph.

### Language Pattern Detection

The **language-lesson** analyzer detects reusable language patterns—such as **generics, closures, and decorators**—and annotates the graph with concept-level explanations. You can see the pattern detection implementation in [`understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts) (lines 52-56).

## Inside the Unified Knowledge Graph

All processed data is merged into a single **Knowledge Graph** stored at [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). The graph supports several node types:

- `file` — a source-code or config file
- `function` — a function or method definition
- `class` — a class definition
- `import` / `export` — dependency edges between files
- `domain` — a business domain such as *payment* or *auth*
- `concept` — a language concept such as *decorator*
- `entity` — a wiki entity extracted from documentation

As described in [`understand-anything-plugin/README.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/README.md) (lines 78-84), **LLM agents**—including logic in [`understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts)—read the structural information together with raw file contents to generate human-readable **summaries, tags, layer assignments, and guided tours**. This hybrid design guarantees that the graph is **deterministic on the structural side** while still capturing the *intent* and *meaning* of the codebase through LLM-generated prose.

## Querying Processed Data Programmatically

The core API—re-exported from [`understand-anything-plugin/packages/core/src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/index.ts)—exposes builders, search engines, and plugins so you can interact with the ingested data directly.

### Run the Full Pipeline from the CLI

```bash
/understand --auto-update

```

### Build the Knowledge Graph in Node.js

```javascript
import { GraphBuilder } from '@understand-anything/core';

const builder = new GraphBuilder();
const graph = await builder.build(); // reads all files, runs Tree-sitter & LLM agents
console.log(graph.nodes.length, graph.edges.length);

```

### Search by Business Domain

```javascript
import { SearchEngine } from '@understand-anything/core';

const engine = new SearchEngine();
const results = await engine.search('payment', { domain: true });

```

### Extract Call-Graph Edges with Tree-sitter

```javascript
import { TreeSitterPlugin } from '@understand-anything/core';

const ts = new TreeSitterPlugin();
await ts.init();
const callGraph = ts.extractCallGraph('src/payment/checkout.ts', sourceCode);

```

## Summary

- Understand Anything processes **any structured text source**, including source code, JSON/YAML configs, Dockerfiles, SQL, markdown wikis, and business-domain descriptors.
- The **Tree-sitter plugin** ([`tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/tree-sitter-plugin.ts)) deterministically extracts functions, classes, imports, and call graphs from code files.
- Dedicated parsers in [`parsers/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/parsers/index.ts) handle non-code configuration files to capture build and infrastructure dependencies.
- LLM-powered agents—**article-analyzer**, **domain-analyzer**, and **language-lesson**—add semantic meaning, business context, and language-concept annotations.
- All data merges into [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) with node types such as `file`, `function`, `domain`, `concept`, and `entity`.
- The core API ([`packages/core/src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/index.ts)) exposes programmatic interfaces to build, search, and traverse the graph.

## Frequently Asked Questions

### Does Understand Anything only process programming language files?

No. While Tree-sitter handles source code, the platform also ingests **JSON, YAML, TOML, Dockerfile, SQL, GraphQL, protobuf, Terraform, Makefile, Shell, and Env** files through its parser plugins. It also processes markdown documentation and wiki content via the article-analyzer agent.

### What node types are stored in the Knowledge Graph?

The graph stores `file`, `function`, `class`, `import`, `export`, `domain`, `concept`, and `entity` nodes. These are assembled into [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) by the `GraphBuilder` orchestrator in [`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts).

### How does Understand Anything extract meaning from documentation?

The **article-analyzer** agent, described in the README, reads markdown-based wikis and extracts **entities, claims, and implicit relationships**. This creates a knowledge graph of ideas that complements the technical graph derived from code.

### Is the Knowledge Graph deterministic?

Yes, on the structural side. Tree-sitter parsing and configuration file parsers always produce the same edges for the same input. LLM agents then provide non-deterministic but rich semantic layers—summaries, tags, and business-domain annotations—that capture intent and meaning.