Core Packages in Lum1104/Understand-Anything: Purpose and Architecture Explained

The @understand-anything/core package is a modular static-analysis engine that transforms source trees into a queryable knowledge graph through specialized sub-modules for parsing, search, persistence, and LLM enrichment.

The Lum1104/Understand-Anything repository turns any codebase into an interactive, knowledge-rich graph. At the center of this system is the @understand-anything/core package, a pipeline of cohesive TypeScript modules that handle everything from AST generation to semantic search. If you want to understand the purpose of each core package in Lum1104/Understand-Anything, this guide maps every key directory and file to its exact responsibility in the analysis pipeline.

Data Model and Schema Validation

Types and Data Model

src/types.ts establishes the canonical schema for the entire knowledge graph. It defines node and edge type enums, the GraphNode and GraphEdge interfaces, the top-level KnowledgeGraph interface, and supporting structures like ThemeConfig that the dashboard consumes. Every other module in the project imports these definitions to ensure structural consistency.

Schema Validation

src/schema.ts enforces structural correctness by validating a generated graph against the JSON schema derived from the type definitions. This step prevents malformed graphs from reaching the UI or persistence layer.

Language Parsing and Plugin Architecture

Plugin Registry

src/plugins/registry.ts and src/plugins/discovery.ts manage the AnalyzerPlugin lifecycle. The registry maps file extensions to their corresponding language extractors, while the discovery utilities scan the environment for available plugins. This design makes it straightforward to add support for new languages without touching the core pipeline.

Tree-Sitter Integration

src/plugins/tree-sitter-plugin.ts wraps the WebAssembly-based web-tree-sitter parser to produce ASTs for every supported language. It exposes a generic parse interface that the language extractors call, isolating parser complexity behind a stable boundary.

Language Extractors

Concrete extractor implementations live under src/plugins/extractors/ and include files such as typescript-extractor.ts, python-extractor.ts, and java-extractor.ts. These modules walk the Tree-Sitter AST and emit a StructuralAnalysis object capturing functions, classes, imports, and other program entities for a given source file.

Building and Refining the Knowledge Graph

Graph Builder

src/analyzer/graph-builder.ts orchestrates the end-to-end pipeline. It reads files from disk, invokes the appropriate language extractors, resolves import relationships, batches nodes and edges, and merges everything into the final KnowledgeGraph. This is the primary workhorse that the public API invokes.

Graph Normalizer

src/analyzer/normalize-graph.ts post-processes the raw graph to produce a clean, stable representation. It performs deduplication, resolves cross-references, and runs layer detection so the dashboard receives a consistent data structure.

Ignore Handling

src/ignore-generator.ts and src/ignore-filter.ts implement an .understandignore style filter. They identify and omit large generated files, binaries, and irrelevant paths before analysis begins, keeping the graph focused and the build fast.

Fingerprinting and Change Detection

src/fingerprint.ts computes a deterministic hash of a file’s contents plus its extracted structural data. The analyzer uses this fingerprint to detect changes and skip unchanged files in subsequent runs. When a file does change, src/change-classifier.ts inspects diffs between successive graph snapshots and labels them—e.g., “added function” or “renamed class”—to drive incremental updates and changelogs.

Search and Semantic Retrieval

src/search.ts provides fast full-text and attribute search over the graph. It pre-builds indexes on node names, tags, and summaries, and exposes a simple search(query) API that the dashboard and CLI can call directly.

src/embedding-search.ts converts node summaries into vector embeddings via an LLM and provides similarity search. This enables semantic queries that match concepts even when the exact keywords differ.

LLM Enrichment and Learning Tours

LLM Analyzer

src/analyzer/llm-analyzer.ts calls a large language model to enrich the graph with higher-level insights. It generates code summaries, documentation blurbs, and suggested learning tours that go beyond what static analysis alone can infer.

Layer and Tour Generation

src/analyzer/layer-detector.ts, src/analyzer/tour-generator.ts, and src/analyzer/language-lesson.ts detect logical groupings (layers) within the codebase and assemble interactive, curriculum-like tours. These tours guide users through the graph in a structured learning flow.

Persistence and Utility Helpers

Persistence Layer

src/persistence/index.ts serializes the generated graph and associated metadata to the .understand-anything/ directory on disk. Fast reloads and incremental updates rely on this storage layer to avoid recomputing the entire graph on every invocation.

Pipeline Utilities

Utility modules such as src/staleness.ts, src/ignore-generator.ts, and src/ignore-filter.ts provide supporting glue for the pipeline. Staleness detection determines whether cached analysis results are still valid, while the ignore utilities keep the file set clean.

Public API Entry Point

src/index.ts exports the clean public API that ties the modules together. Consumers such as the dashboard or CLI import functions like buildGraph, search, and llmAnalyze to run the full pipeline without dealing with internal plumbing.

import { buildGraph } from '@understand-anything/core';

// Build a complete knowledge graph for a project folder
const graph = await buildGraph({
  root: '/path/to/project',
  ignoreFile: '.understandignore',
});

// Perform a keyword search
const results = await search(graph, 'authentication');

// Run an LLM-enriched analysis (requires an LLM endpoint)
await llmAnalyze(graph, { model: 'gpt-4o' });

Summary

The Lum1104/Understand-Anything core package is organized into clear functional layers:

Frequently Asked Questions

What is the main entry point of the Understand Anything core package?

src/index.ts serves as the public API boundary. It exports high-level functions such as buildGraph and search so that the dashboard and CLI can invoke the entire pipeline without importing individual internal modules.

How does the core package avoid re-analyzing unchanged files?

src/fingerprint.ts computes a deterministic hash of each file’s contents and its extracted structural data. The pipeline compares this fingerprint across runs and skips any file whose hash has not changed, while src/change-classifier.ts categorizes actual modifications for incremental updates.

Which file handles the actual AST parsing for different programming languages?

src/plugins/tree-sitter-plugin.ts wraps the WebAssembly web-tree-sitter parser and exposes a generic parse interface. Individual languages then use their own extractor files—such as src/plugins/extractors/typescript-extractor.ts and src/plugins/extractors/python-extractor.ts—to traverse the resulting AST.

Can the knowledge graph be queried semantically beyond keyword matching?

Yes. src/search.ts provides traditional full-text search over names, tags, and summaries. For semantic queries, src/embedding-search.ts converts node summaries into vector embeddings via an LLM and performs similarity search, enabling concept-based retrieval that does not rely on exact keyword matches.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →