What Kind of Data Does Understand Anything Process? Sources, Node Types, and Code Examples
Understand Anything ingests any structured text source—including source code, configuration files, documentation, and business-domain descriptors—to build a unified Knowledge Graph that mixes deterministic syntax extraction with LLM-generated semantic meaning.
The open-source repository Lum1104/Understand-Anything is designed to transform any project artifact that can be represented as text into a rich, queryable graph. What kind of data Understand Anything processes spans from raw source files parsed by Tree-sitter to markdown wikis analyzed by LLM agents. All extracted nodes and edges are merged into a single Knowledge Graph at .understand-anything/knowledge-graph.json, which powers search, diff-impact analysis, and guided tours.
Source Code Files Parsed by Tree-sitter
At the core of the pipeline, the Tree-sitter plugin scans every file in a project and parses it into a concrete syntax tree. In understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts (lines 21-30), the plugin extracts functions, classes, imports/exports, and call-graph edges regardless of language. This structural layer is fully deterministic: the same code always yields the same nodes and relationships.
Configuration and Build Artifacts
Beyond source code, Understand Anything processes infrastructure and build files through dedicated parsers. The index at understand-anything-plugin/packages/core/src/plugins/parsers/index.ts (lines 4-17) registers handlers for JSON, YAML, TOML, Dockerfile, SQL, GraphQL, protobuf, Terraform, Makefile, Shell, and Env files. These parsers turn configuration data into structured nodes, allowing the graph to capture configuration-driven dependencies alongside code relationships.
Documentation and Knowledge Bases
The pipeline treats documentation as first-class data. According to understand-anything-plugin/README.md (lines 65-68), markdown-based wikis—such as “Karpathy-pattern” LLM wikis—are fed to the article-analyzer agent. This agent extracts entities, claims, and implicit relationships to build a knowledge graph of ideas that sits alongside the technical graph.
Semantic Layers: Business Domains and Language Concepts
After structural extraction, specialized analyzers annotate the graph with higher-level meaning.
Business-Domain Tagging
The domain-analyzer agent—centralized in understand-anything-plugin/packages/core/src/analyzer/domain-analyzer.ts and orchestrated within understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts (lines 89-97)—tags code with domains, flows, and process steps. This turns technical artifacts into business-oriented concepts that are stored as domain nodes in the final graph.
Language Pattern Detection
The language-lesson analyzer detects reusable language patterns—such as generics, closures, and decorators—and annotates the graph with concept-level explanations. You can see the pattern detection implementation in understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts (lines 52-56).
Inside the Unified Knowledge Graph
All processed data is merged into a single Knowledge Graph stored at .understand-anything/knowledge-graph.json. The graph supports several node types:
file— a source-code or config filefunction— a function or method definitionclass— a class definitionimport/export— dependency edges between filesdomain— a business domain such as payment or authconcept— a language concept such as decoratorentity— a wiki entity extracted from documentation
As described in understand-anything-plugin/README.md (lines 78-84), LLM agents—including logic in understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts—read the structural information together with raw file contents to generate human-readable summaries, tags, layer assignments, and guided tours. This hybrid design guarantees that the graph is deterministic on the structural side while still capturing the intent and meaning of the codebase through LLM-generated prose.
Querying Processed Data Programmatically
The core API—re-exported from understand-anything-plugin/packages/core/src/index.ts—exposes builders, search engines, and plugins so you can interact with the ingested data directly.
Run the Full Pipeline from the CLI
/understand --auto-update
Build the Knowledge Graph in Node.js
import { GraphBuilder } from '@understand-anything/core';
const builder = new GraphBuilder();
const graph = await builder.build(); // reads all files, runs Tree-sitter & LLM agents
console.log(graph.nodes.length, graph.edges.length);
Search by Business Domain
import { SearchEngine } from '@understand-anything/core';
const engine = new SearchEngine();
const results = await engine.search('payment', { domain: true });
Extract Call-Graph Edges with Tree-sitter
import { TreeSitterPlugin } from '@understand-anything/core';
const ts = new TreeSitterPlugin();
await ts.init();
const callGraph = ts.extractCallGraph('src/payment/checkout.ts', sourceCode);
Summary
- Understand Anything processes any structured text source, including source code, JSON/YAML configs, Dockerfiles, SQL, markdown wikis, and business-domain descriptors.
- The Tree-sitter plugin (
tree-sitter-plugin.ts) deterministically extracts functions, classes, imports, and call graphs from code files. - Dedicated parsers in
parsers/index.tshandle non-code configuration files to capture build and infrastructure dependencies. - LLM-powered agents—article-analyzer, domain-analyzer, and language-lesson—add semantic meaning, business context, and language-concept annotations.
- All data merges into
.understand-anything/knowledge-graph.jsonwith node types such asfile,function,domain,concept, andentity. - The core API (
packages/core/src/index.ts) exposes programmatic interfaces to build, search, and traverse the graph.
Frequently Asked Questions
Does Understand Anything only process programming language files?
No. While Tree-sitter handles source code, the platform also ingests JSON, YAML, TOML, Dockerfile, SQL, GraphQL, protobuf, Terraform, Makefile, Shell, and Env files through its parser plugins. It also processes markdown documentation and wiki content via the article-analyzer agent.
What node types are stored in the Knowledge Graph?
The graph stores file, function, class, import, export, domain, concept, and entity nodes. These are assembled into .understand-anything/knowledge-graph.json by the GraphBuilder orchestrator in understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts.
How does Understand Anything extract meaning from documentation?
The article-analyzer agent, described in the README, reads markdown-based wikis and extracts entities, claims, and implicit relationships. This creates a knowledge graph of ideas that complements the technical graph derived from code.
Is the Knowledge Graph deterministic?
Yes, on the structural side. Tree-sitter parsing and configuration file parsers always produce the same edges for the same input. LLM agents then provide non-deterministic but rich semantic layers—summaries, tags, and business-domain annotations—that capture intent and meaning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →