How Language Extractors Work for TypeScript, Python, and Go in Understand-Anything

Understand-Anything uses tree-sitter based language extractors that parse TypeScript, Python, and Go source files into ASTs, then walk them to generate symbol information that gets assembled into a unified knowledge graph.

The Egonex-AI/Understand-Anything repository implements a language-agnostic knowledge graph builder that processes source code through specialized extractors. These components parse code into abstract syntax trees and extract semantic symbols using the tree-sitter parsing library. The extraction pipeline supports multiple languages including TypeScript, Python, and Go through a pluggable architecture defined in the core package.

The Language Extraction Pipeline

The extraction process follows six distinct stages that transform raw source files into structured graph nodes.

File Discovery and Language Detection

The LanguageRegistry in packages/core/src/languages/language-registry.ts maps file extensions to language IDs. For example, .ts maps to typescript, .py to python, and .go to go.

Tree-Sitter Parser Loading

The TreeSitterPlugin defined in packages/core/src/plugins/tree-sitter-plugin.ts loads the appropriate WASM parser for each language. It creates TreeSitterLanguage instances using modules like tree-sitter-typescript, tree-sitter-python, and tree-sitter-go.

AST Generation

The parser processes file content and returns a syntax tree representation that extractors can traverse.

Symbol Extraction

Concrete extractor implementations walk the AST using tree-sitter's query API to identify classes, functions, imports, and type definitions. Each language has specific patterns it searches for.

Graph Enrichment

The GraphBuilder in packages/core/src/analyzer/graph-builder.ts normalizes extracted symbols using the schema defined in packages/core/src/schema.ts. It assigns language-specific concepts like "class" or "function" and links them into the project-wide graph.

Output Persistence

The final knowledge graph is serialized to .understand-anything/knowledge-graph.json for consumption by the dashboard.

Language-Specific Extractor Implementations

Each extractor implements the LanguageExtractor interface from packages/core/src/plugins/extractors/base-extractor.ts.

TypeScript Extractor

The TypeScript extractor in typescript-extractor.ts handles both TypeScript and JavaScript files. It extracts classes, interfaces, functions, and type aliases by traversing the TypeScript AST. The extractor treats JavaScript as a sibling language, ensuring compatibility with mixed JS/TS projects.

Python Extractor

Located in python-extractor.ts, this extractor identifies module-level and nested definitions using queries for def, class, and import statements. It captures both top-level functions and classes as well as nested definitions within methods.

Go Extractor

The Go extractor in go-extractor.ts processes package declarations, func, type, var, and import statements. It walks the Go AST to capture the module structure and symbol relationships specific to Go's package system.

Shared Extraction Interface

All extractors share a common contract defined in the base extractor:

  • languageIds: An array of language IDs the extractor supports (e.g., ['typescript', 'tsx'])
  • extract(source: string, filePath: string): Parses the source, runs language-specific queries, and returns an array of SymbolInfo objects
  • getLanguageConcepts(node: TreeSitterNode, language: string): Maps raw syntax nodes to high-level concepts used by the UI

Because implementations depend only on tree-sitter WASM modules, they run identically in Node.js and browser environments.

Running the Extractors

You can invoke extractors manually or through the automated Analyzer pipeline.

// Manual extraction example
import { TypeScriptExtractor } from '@understand-anything/core/plugins/extractors/typescript-extractor';
import { readFile } from 'fs/promises';

async function extractTs(file: string) {
  const src = await readFile(file, 'utf8');
  const extractor = new TypeScriptExtractor();
  const symbols = extractor.extract(src, file);
  console.log(symbols);
}

extractTs('src/app.ts');
// Automated analysis using the Analyzer
import { Analyzer } from '@understand-anything/core/analyzer';

const project = {
  root: '/my/project',
  files: ['src/main.py', 'src/util.go', 'src/component.tsx'],
};

const analyzer = new Analyzer();
const graph = await analyzer.analyze(project);
console.log(graph.nodes.filter(n => n.language === 'python'));

Key Source Files

Summary

  • Language extractors in Understand-Anything implement a common interface defined in base-extractor.ts
  • The LanguageRegistry maps file extensions to appropriate extractors automatically
  • Tree-sitter WASM parsers generate ASTs for TypeScript, Python, and Go source files
  • Each extractor walks the AST to emit SymbolInfo objects representing code symbols
  • The GraphBuilder consolidates extracted symbols into a unified knowledge graph stored at .understand-anything/knowledge-graph.json
  • The architecture supports both Node.js and browser execution environments

Frequently Asked Questions

How does the LanguageRegistry determine which extractor to use?

The LanguageRegistry in packages/core/src/languages/language-registry.ts maintains a mapping of file extensions to language IDs. When a file is processed, the registry examines the extension (e.g., .ts, .py, .go) and returns the corresponding language ID, which the Analyzer uses to instantiate the appropriate extractor implementation.

What parsing library powers the language extractors?

The extractors rely on tree-sitter, a parser generator tool and incremental parsing library. The TreeSitterPlugin in packages/core/src/plugins/tree-sitter-plugin.ts loads language-specific WASM parsers such as tree-sitter-typescript, tree-sitter-python, and tree-sitter-go to parse source code into ASTs that extractors can traverse.

Can extractors run outside of Node.js environments?

Yes, because the extractors depend only on tree-sitter WASM modules rather than native bindings, they execute identically in Node.js and browser environments. This design allows the Understand-Anything dashboard to perform client-side analysis without requiring a backend server.

How are extracted symbols structured for the knowledge graph?

Each extractor returns SymbolInfo objects from the extract(source, filePath) method. The GraphBuilder in packages/core/src/analyzer/graph-builder.ts processes these objects using the schema defined in packages/core/src/schema.ts, normalizing language-specific concepts like "class" or "function" into a unified graph format stored in .understand-anything/knowledge-graph.json.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →