What Data Sources Does Understand-Anything Support? A Complete Guide
Understand-Anything ingests code repositories across 30+ languages via Tree-sitter, Markdown-centric knowledge bases such as Obsidian and Logseq vaults, and general documentation files, unifying all of them into a single knowledge graph for interactive exploration.
The open-source Lum1104/Understand-Anything project transforms disparate inputs into a queryable knowledge graph that powers its dashboard, search, tours, and LLM-assisted explanations. Whether you are scanning a codebase, a personal wiki, or a folder of design documents, the tool maps entities and relationships into a shared schema. This guide covers every supported source and how each is processed based on the source code.
Code Repositories and Source Files
Understand-Anything accepts any repository of source files. Its static analysis layer uses Tree-sitter parsers for every language that has a Tree-sitter grammar, which includes more than 30 languages such as JavaScript/TypeScript, Python, Go, Java, Kotlin, Rust, C/C++, and Ruby. The parsers deterministically extract imports, definitions, calls, and other structural facts that become nodes and edges in the graph.
In understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts, the system performs language detection and drives the Tree-sitter integration. Files with supported grammars receive full parsing, while unsupported files fall back to a content-hash fingerprint that indexes their raw text. This fallback ensures that every file appears in the graph even when it cannot be structurally parsed.
The core type definitions in understand-anything-plugin/packages/core/src/types.ts define node kinds such as file, function, and class. After parsing, an LLM layer adds summaries, tags, and layer assignments to the graph.
Markdown Knowledge Bases
The /understand-knowledge skill handles any collection of markdown files. According to the design spec in docs/superpowers/specs/2026-04-09-understand-knowledge-design.md, the skill scans the target directory, builds a manifest of markdown files, detects the layout, and parses wiki-style links, tags, and front-matter.
Supported Knowledge Base Formats
The parser automatically detects the following markdown-centric structures:
- Obsidian vaults
- Logseq graphs
- Dendron workspaces
- Foam notebooks
- Karpathy-style LLM wikis (three-layer format with raw sources, markdown, and a schema file)
- Zettelkasten piles and plain markdown directories
The resulting entities, claims, and topics become typed nodes such as article, entity, topic, claim, and source. Plugin registration for these parsers is coordinated through understand-anything-plugin/packages/core/src/plugins/registry.ts.
General Documentation and Mixed Inputs
Beyond code and structured knowledge bases, Understand-Anything ingests general documentation. Any non-code text file—such as README files, design docs, *.md, or *.txt—is handled by the generic file-scanner in the project-scanner agent.
These files receive the same fingerprint-only fallback used for unsupported code files. They appear as document nodes in the graph and connect to other nodes via references edges. When source text is unavailable for rendering, the dashboard references sourceUnavailable strings in its locale files.
How Understand-Anything Processes Each Source
All inputs converge into a unified knowledge graph stored at .understand-anything/knowledge-graph.json. The pipeline follows three consistent stages:
- Discovery – The scanner identifies relevant files based on the entry point (
/understandor/understand-knowledge). - Parsing – Tree-sitter handles supported code; the markdown skill handles knowledge bases; generic text files receive hash-based indexing.
- Enrichment – An LLM layer generates summaries, tags, and structural assignments for the parsed nodes.
Because the same graph schema in packages/core/src/types.ts powers the dashboard, search, and tours, you can navigate across code, docs, and knowledge bases seamlessly.
Quick Start Commands
You can analyze any supported source using these entry points:
# Analyze a code repository (auto-discovers code and adjacent docs)
/understand
# Analyze a code repository limited to a subdirectory
/understand src/
# Analyze a markdown knowledge base
/understand-knowledge /path/to/wiki
# Launch the interactive dashboard
/understand-dashboard
For example, to process an Obsidian vault:
/understand-knowledge ~/Documents/ObsidianVault
To mix code and documentation in a single graph:
/understand
# Picks up *.js, *.ts, *.md, *.txt, and builds a unified graph.
Summary
- Code repositories are parsed with Tree-sitter for 30+ languages, with a content-hash fallback for unsupported files, as implemented in
packages/core/src/plugins/tree-sitter-plugin.ts. - Knowledge bases include Obsidian, Logseq, Dendron, Foam, Karpathy-style wikis, and Zettelkasten, parsed by the
/understand-knowledgeskill. - General text documents are indexed as
documentnodes and linked viareferencesedges in the graph. - All sources feed into the shared type system defined in
packages/core/src/types.tsand are explorable through the interactive dashboard.
Frequently Asked Questions
Does Understand-Anything support languages without a Tree-sitter grammar?
Yes. If a file’s language lacks a Tree-sitter grammar, the scanner falls back to a content-hash fingerprint. The file is indexed by its raw text and appears in the graph, but it does not receive structural parsing for imports or definitions. This behavior is managed by the tree-sitter plugin.
Can I analyze an Obsidian vault and a codebase at the same time?
Yes. Running /understand in a directory that contains both source code and markdown documentation will auto-discover both file types and produce a single unified graph. You can also run /understand-knowledge on a vault separately if you only want the knowledge-base view.
Where is the generated knowledge graph stored?
After parsing, the graph is stored at .understand-anything/knowledge-graph.json by default. You can explore it by running /understand-dashboard, which opens the browser UI and allows you to inspect nodes, read LLM-generated summaries, and view source links.
What node types are created from markdown knowledge bases?
The knowledge-base parser creates typed nodes including article, entity, topic, claim, and source. These nodes capture wiki links, tags, and front-matter extracted from your markdown files, as defined in the core types and the knowledge-base design spec.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →