# What Data Sources Does Understand-Anything Support? A Complete Guide

> Discover the data sources Lum1104/Understand-Anything supports including 30+ code languages, Obsidian, Logseq, and docs. Explore your unified knowledge graph.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-07

---

**Understand-Anything ingests code repositories across 30+ languages via Tree-sitter, Markdown-centric knowledge bases such as Obsidian and Logseq vaults, and general documentation files, unifying all of them into a single knowledge graph for interactive exploration.**

The open-source `Lum1104/Understand-Anything` project transforms disparate inputs into a queryable knowledge graph that powers its dashboard, search, tours, and LLM-assisted explanations. Whether you are scanning a codebase, a personal wiki, or a folder of design documents, the tool maps entities and relationships into a shared schema. This guide covers every supported source and how each is processed based on the source code.

## Code Repositories and Source Files

Understand-Anything accepts any repository of source files. Its static analysis layer uses **Tree-sitter** parsers for every language that has a Tree-sitter grammar, which includes more than 30 languages such as JavaScript/TypeScript, Python, Go, Java, Kotlin, Rust, C/C++, and Ruby. The parsers deterministically extract imports, definitions, calls, and other structural facts that become nodes and edges in the graph.

In [`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts), the system performs language detection and drives the Tree-sitter integration. Files with supported grammars receive full parsing, while unsupported files fall back to a **content-hash fingerprint** that indexes their raw text. This fallback ensures that every file appears in the graph even when it cannot be structurally parsed.

The core type definitions in [`understand-anything-plugin/packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/types.ts) define node kinds such as `file`, `function`, and `class`. After parsing, an LLM layer adds summaries, tags, and layer assignments to the graph.

## Markdown Knowledge Bases

The `/understand-knowledge` skill handles any collection of markdown files. According to the design spec in [`docs/superpowers/specs/2026-04-09-understand-knowledge-design.md`](https://github.com/Lum1104/Understand-Anything/blob/main/docs/superpowers/specs/2026-04-09-understand-knowledge-design.md), the skill scans the target directory, builds a manifest of markdown files, detects the layout, and parses wiki-style links, tags, and front-matter.

### Supported Knowledge Base Formats

The parser automatically detects the following markdown-centric structures:

- **Obsidian vaults**
- **Logseq graphs**
- **Dendron workspaces**
- **Foam notebooks**
- **Karpathy-style LLM wikis** (three-layer format with raw sources, markdown, and a schema file)
- **Zettelkasten** piles and plain markdown directories

The resulting entities, claims, and topics become typed nodes such as `article`, `entity`, `topic`, `claim`, and `source`. Plugin registration for these parsers is coordinated through [`understand-anything-plugin/packages/core/src/plugins/registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/registry.ts).

## General Documentation and Mixed Inputs

Beyond code and structured knowledge bases, Understand-Anything ingests general documentation. Any non-code text file—such as README files, design docs, `*.md`, or `*.txt`—is handled by the generic file-scanner in the **project-scanner** agent.

These files receive the same fingerprint-only fallback used for unsupported code files. They appear as `document` nodes in the graph and connect to other nodes via `references` edges. When source text is unavailable for rendering, the dashboard references `sourceUnavailable` strings in its locale files.

## How Understand-Anything Processes Each Source

All inputs converge into a unified knowledge graph stored at [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). The pipeline follows three consistent stages:

1. **Discovery** – The scanner identifies relevant files based on the entry point (`/understand` or `/understand-knowledge`).
2. **Parsing** – Tree-sitter handles supported code; the markdown skill handles knowledge bases; generic text files receive hash-based indexing.
3. **Enrichment** – An LLM layer generates summaries, tags, and structural assignments for the parsed nodes.

Because the same graph schema in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts) powers the dashboard, search, and tours, you can navigate across code, docs, and knowledge bases seamlessly.

## Quick Start Commands

You can analyze any supported source using these entry points:

```bash

# Analyze a code repository (auto-discovers code and adjacent docs)

/understand

# Analyze a code repository limited to a subdirectory

/understand src/

# Analyze a markdown knowledge base

/understand-knowledge /path/to/wiki

# Launch the interactive dashboard

/understand-dashboard

```

For example, to process an Obsidian vault:

```bash
/understand-knowledge ~/Documents/ObsidianVault

```

To mix code and documentation in a single graph:

```bash
/understand

# Picks up *.js, *.ts, *.md, *.txt, and builds a unified graph.

```

## Summary

- **Code repositories** are parsed with Tree-sitter for 30+ languages, with a content-hash fallback for unsupported files, as implemented in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts).
- **Knowledge bases** include Obsidian, Logseq, Dendron, Foam, Karpathy-style wikis, and Zettelkasten, parsed by the `/understand-knowledge` skill.
- **General text documents** are indexed as `document` nodes and linked via `references` edges in the graph.
- All sources feed into the shared type system defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts) and are explorable through the interactive dashboard.

## Frequently Asked Questions

### Does Understand-Anything support languages without a Tree-sitter grammar?

Yes. If a file’s language lacks a Tree-sitter grammar, the scanner falls back to a content-hash fingerprint. The file is indexed by its raw text and appears in the graph, but it does not receive structural parsing for imports or definitions. This behavior is managed by the tree-sitter plugin.

### Can I analyze an Obsidian vault and a codebase at the same time?

Yes. Running `/understand` in a directory that contains both source code and markdown documentation will auto-discover both file types and produce a single unified graph. You can also run `/understand-knowledge` on a vault separately if you only want the knowledge-base view.

### Where is the generated knowledge graph stored?

After parsing, the graph is stored at [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) by default. You can explore it by running `/understand-dashboard`, which opens the browser UI and allows you to inspect nodes, read LLM-generated summaries, and view source links.

### What node types are created from markdown knowledge bases?

The knowledge-base parser creates typed nodes including `article`, `entity`, `topic`, `claim`, and `source`. These nodes capture wiki links, tags, and front-matter extracted from your markdown files, as defined in the core types and the knowledge-base design spec.