# Core Packages in Lum1104/Understand-Anything: Purpose and Architecture Explained

> Explore the core packages in Lum1104/Understand-Anything. Discover how this static-analysis engine builds a queryable knowledge graph from your source code for parsing, search, persistence, and LLM enrichment.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: architecture
- Published: 2026-06-07

---

**The `@understand-anything/core` package is a modular static-analysis engine that transforms source trees into a queryable knowledge graph through specialized sub-modules for parsing, search, persistence, and LLM enrichment.**

The `Lum1104/Understand-Anything` repository turns any codebase into an interactive, knowledge-rich graph. At the center of this system is the `@understand-anything/core` package, a pipeline of cohesive TypeScript modules that handle everything from AST generation to semantic search. If you want to understand the purpose of each core package in `Lum1104/Understand-Anything`, this guide maps every key directory and file to its exact responsibility in the analysis pipeline.

## Data Model and Schema Validation

### Types and Data Model

[`src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/types.ts) establishes the canonical schema for the entire knowledge graph. It defines **node and edge type enums**, the `GraphNode` and `GraphEdge` interfaces, the top-level `KnowledgeGraph` interface, and supporting structures like `ThemeConfig` that the dashboard consumes. Every other module in the project imports these definitions to ensure structural consistency.

### Schema Validation

[`src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/schema.ts) enforces structural correctness by validating a generated graph against the JSON schema derived from the type definitions. This step prevents malformed graphs from reaching the UI or persistence layer.

## Language Parsing and Plugin Architecture

### Plugin Registry

[`src/plugins/registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/registry.ts) and [`src/plugins/discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/discovery.ts) manage the **AnalyzerPlugin** lifecycle. The registry maps file extensions to their corresponding language extractors, while the discovery utilities scan the environment for available plugins. This design makes it straightforward to add support for new languages without touching the core pipeline.

### Tree-Sitter Integration

[`src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/tree-sitter-plugin.ts) wraps the WebAssembly-based `web-tree-sitter` parser to produce ASTs for every supported language. It exposes a generic `parse` interface that the language extractors call, isolating parser complexity behind a stable boundary.

### Language Extractors

Concrete extractor implementations live under `src/plugins/extractors/` and include files such as [`typescript-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/typescript-extractor.ts), [`python-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/python-extractor.ts), and [`java-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/java-extractor.ts). These modules walk the Tree-Sitter AST and emit a `StructuralAnalysis` object capturing functions, classes, imports, and other program entities for a given source file.

## Building and Refining the Knowledge Graph

### Graph Builder

[`src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/graph-builder.ts) orchestrates the end-to-end pipeline. It reads files from disk, invokes the appropriate language extractors, resolves import relationships, batches nodes and edges, and merges everything into the final `KnowledgeGraph`. This is the primary workhorse that the public API invokes.

### Graph Normalizer

[`src/analyzer/normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/normalize-graph.ts) post-processes the raw graph to produce a clean, stable representation. It performs deduplication, resolves cross-references, and runs layer detection so the dashboard receives a consistent data structure.

### Ignore Handling

[`src/ignore-generator.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/ignore-generator.ts) and [`src/ignore-filter.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/ignore-filter.ts) implement an `.understandignore` style filter. They identify and omit large generated files, binaries, and irrelevant paths before analysis begins, keeping the graph focused and the build fast.

### Fingerprinting and Change Detection

[`src/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/fingerprint.ts) computes a deterministic hash of a file’s contents plus its extracted structural data. The analyzer uses this fingerprint to detect changes and skip unchanged files in subsequent runs. When a file does change, [`src/change-classifier.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/change-classifier.ts) inspects diffs between successive graph snapshots and labels them—e.g., “added function” or “renamed class”—to drive incremental updates and changelogs.

## Search and Semantic Retrieval

### Full-Text Search

[`src/search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/search.ts) provides fast full-text and attribute search over the graph. It pre-builds indexes on node names, tags, and summaries, and exposes a simple `search(query)` API that the dashboard and CLI can call directly.

### Embedding-Based Search

[`src/embedding-search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/embedding-search.ts) converts node summaries into vector embeddings via an LLM and provides similarity search. This enables semantic queries that match concepts even when the exact keywords differ.

## LLM Enrichment and Learning Tours

### LLM Analyzer

[`src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/llm-analyzer.ts) calls a large language model to enrich the graph with higher-level insights. It generates code summaries, documentation blurbs, and suggested learning tours that go beyond what static analysis alone can infer.

### Layer and Tour Generation

[`src/analyzer/layer-detector.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/layer-detector.ts), [`src/analyzer/tour-generator.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/tour-generator.ts), and [`src/analyzer/language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/language-lesson.ts) detect logical groupings (layers) within the codebase and assemble interactive, curriculum-like tours. These tours guide users through the graph in a structured learning flow.

## Persistence and Utility Helpers

### Persistence Layer

[`src/persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/persistence/index.ts) serializes the generated graph and associated metadata to the `.understand-anything/` directory on disk. Fast reloads and incremental updates rely on this storage layer to avoid recomputing the entire graph on every invocation.

### Pipeline Utilities

Utility modules such as [`src/staleness.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/staleness.ts), [`src/ignore-generator.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/ignore-generator.ts), and [`src/ignore-filter.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/ignore-filter.ts) provide supporting glue for the pipeline. Staleness detection determines whether cached analysis results are still valid, while the ignore utilities keep the file set clean.

## Public API Entry Point

[`src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/index.ts) exports the clean public API that ties the modules together. Consumers such as the dashboard or CLI import functions like `buildGraph`, `search`, and `llmAnalyze` to run the full pipeline without dealing with internal plumbing.

```ts
import { buildGraph } from '@understand-anything/core';

// Build a complete knowledge graph for a project folder
const graph = await buildGraph({
  root: '/path/to/project',
  ignoreFile: '.understandignore',
});

// Perform a keyword search
const results = await search(graph, 'authentication');

// Run an LLM-enriched analysis (requires an LLM endpoint)
await llmAnalyze(graph, { model: 'gpt-4o' });

```

## Summary

The `Lum1104/Understand-Anything` core package is organized into clear functional layers:

- **Data model**: [`src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/types.ts) and [`src/schema.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/schema.ts) define and validate the graph schema.
- **Parsing**: [`src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/tree-sitter-plugin.ts) and language extractors turn source files into ASTs and structured analysis objects.
- **Construction**: [`src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/graph-builder.ts) and [`src/analyzer/normalize-graph.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/normalize-graph.ts) assemble and clean the final graph.
- **Incremental analysis**: [`src/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/fingerprint.ts) and [`src/change-classifier.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/change-classifier.ts) detect and classify changes to avoid redundant work.
- **Search**: [`src/search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/search.ts) and [`src/embedding-search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/embedding-search.ts) power keyword and semantic queries.
- **Enrichment**: [`src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/analyzer/llm-analyzer.ts) and tour generators add AI-driven insights and guided learning.
- **Persistence**: [`src/persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/persistence/index.ts) stores graphs for fast reloads.

## Frequently Asked Questions

### What is the main entry point of the Understand Anything core package?

[`src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/index.ts) serves as the public API boundary. It exports high-level functions such as `buildGraph` and `search` so that the dashboard and CLI can invoke the entire pipeline without importing individual internal modules.

### How does the core package avoid re-analyzing unchanged files?

[`src/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/fingerprint.ts) computes a deterministic hash of each file’s contents and its extracted structural data. The pipeline compares this fingerprint across runs and skips any file whose hash has not changed, while [`src/change-classifier.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/change-classifier.ts) categorizes actual modifications for incremental updates.

### Which file handles the actual AST parsing for different programming languages?

[`src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/tree-sitter-plugin.ts) wraps the WebAssembly `web-tree-sitter` parser and exposes a generic `parse` interface. Individual languages then use their own extractor files—such as [`src/plugins/extractors/typescript-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/extractors/typescript-extractor.ts) and [`src/plugins/extractors/python-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/plugins/extractors/python-extractor.ts)—to traverse the resulting AST.

### Can the knowledge graph be queried semantically beyond keyword matching?

Yes. [`src/search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/search.ts) provides traditional full-text search over names, tags, and summaries. For semantic queries, [`src/embedding-search.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/src/embedding-search.ts) converts node summaries into vector embeddings via an LLM and performs similarity search, enabling concept-based retrieval that does not rely on exact keyword matches.