How Language-Specific Extractors Parse Code Structure into Nodes and Edges in Understand Anything

Language-specific extractors in Egonex-AI/Understand-Anything convert source code into structural graphs and call graphs by using tree-sitter to parse ASTs, then walking the tree to create node entries for functions and classes, and edge entries for function calls.

Understand Anything uses language-specific extractors to transform raw source code into a navigable knowledge graph. These extractors implement the LanguageExtractor interface defined in src/plugins/extractors/types.ts, enabling a unified pipeline that parses code structure into nodes and edges regardless of the programming language. Each extractor performs a two-phase analysis: first building a structural graph of definitions, then constructing a call graph of relationships.

The Extraction Pipeline: From Source to Graph

Every extractor follows a standardized six-step workflow to convert source files into graph data. The process is consistent across all supported languages, differing only in the specific tree-sitter node types each extractor recognizes.

Step Action Key Components
1. Parse Feed source code to tree-sitter to generate a language-specific AST. treeSitterParser.parse(code) in src/plugins/tree-sitter-plugin.ts
2. Walk Iterate over top-level AST children to identify constructs. Root node traversal loop in extractStructure
3. Build Nodes Create StructuralAnalysis entries for functions, classes, enums, and imports. extractFunction, extractStruct, extractEnum, extractTrait
4. Record Edges Traverse the AST again to map caller-callee relationships. extractCallGraph → walkForCalls → extractCalleeName
5. Post-Process Merge auxiliary data (e.g., methods to their owning structs). methodsByType map processing
6. Return Output structured data for the knowledge graph builder. Returns { functions, classes, imports, exports } and CallGraphEntry[]

How Nodes Are Built: Structural Analysis

The first phase extracts nodes representing code entities. Starting at the root of the tree-sitter AST, the extractor iterates through child nodes and delegates to specialized helper functions based on node type.

Walking the AST

In src/plugins/extractors/rust-extractor.ts, the extractStructure method (lines 95-143) implements the traversal:

for (let i = 0; i < rootNode.childCount; i++) {
  const node = rootNode.child(i);
  // Delegation based on node type
}

This pattern appears in every language extractor. The traverse utility in src/plugins/extractors/base-extractor.ts provides a generic depth-first visitor that individual extractors extend for language-specific node types.

Creating Structural Entries

For each identified construct, the extractor populates a StructuralAnalysis object with metadata:

  • Functions: Captured via extractFunction, collecting line ranges, parameter names via extractParams, and return types via extractReturnType
  • Classes/Structs: Extracted using extractStruct or extractEnum, recording visibility via isPublic(node) checks for visibility_modifier children starting with pub
  • Imports: Processed by extractUseDeclaration, with extractScopedPath flattening scoped identifiers like std::collections::HashMap
  • Exports: Public items are filtered and added to the exports list based on visibility checks

How Edges Are Tracked: Call Graph Construction

After structural nodes are established, the extractor performs a second pass to identify edges representing relationships between entities.

The Two-Pass Strategy

The extractCallGraph method runs separately from extractStructure to maintain clean separation between definition discovery and relationship mapping. According to the Rust extractor implementation (lines 150-225), this method maintains a stack tracking the current function context.

Resolving Caller-Callee Relationships

When the walker encounters a call_expression node, it extracts the callee identifier using extractCalleeName and creates a CallGraphEntry containing:

  • caller: The current function from the stack
  • callee: The resolved identifier (plain, field expression, or scoped)
  • line: The source line number

This approach captures function calls, method invocations, and cross-module references regardless of whether the target is defined in the same file.

Language-Specific Implementation: The Rust Extractor Example

The Rust extractor in src/plugins/extractors/rust-extractor.ts demonstrates how language-specific semantics integrate into the generic pipeline.

Visibility and Exports

The isPublic(node) helper checks for visibility_modifier children. Only items marked with pub (or variant modifiers like pub(crate)) are added to the exports array, ensuring the knowledge graph accurately reflects the public API surface.

Method Association via Impl Blocks

Rust requires special handling for impl blocks. The extractor stores method names in a methodsByType map (lines 101-108) during the first pass. After processing all top-level items, it merges these methods into their corresponding struct or enum entries (lines 135-142), creating complete class definitions that include both declared fields and implemented methods.

Practical Usage Example

import { RustExtractor } from "./packages/core/src/plugins/extractors/rust-extractor.js";
import { TreeSitterParser } from "./packages/core/src/plugins/tree-sitter-plugin.js";

const code = `
pub struct Point { x: i32, y: i32 }

impl Point {
    pub fn distance(&self, other: &Point) -> f64 {
        ((self.x - other.x).pow(2) + (self.y - other.y).pow(2)).sqrt()
    }
}

fn main() {
    let a = Point { x: 0, y: 0 };
    let b = Point { x: 3, y: 4 };
    println!("{}", a.distance(&b));
}
`;

const parser = new TreeSitterParser("rust");
const tree = parser.parse(code);
const extractor = new RustExtractor();

const structure = extractor.extractStructure(tree.rootNode);
const callGraph = extractor.extractCallGraph(tree.rootNode);

console.log(structure.classes);  // [{ name: "Point", methods: ["distance"], ... }]
console.log(callGraph);          // [{ caller: "main", callee: "distance", ... }]

Shared Utilities and Base Classes

While each language implements its own extractor, shared utilities in src/plugins/extractors/base-extractor.ts prevent code duplication:

  • traverse: Generic depth-first AST visitor
  • findChild / findChildren: Locate specific node types within the tree
  • getStringValue: Unquotes string literals for languages like Python that expose them as fragments

These utilities allow the Python, TypeScript, Go, and Java extractors to follow the same architectural patterns as the Rust implementation, varying only in the specific tree-sitter node types they handle.

Summary

  • Language-specific extractors implement the LanguageExtractor interface to provide a unified API for parsing different programming languages.
  • Two-phase extraction separates structural analysis (extractStructure) from relationship mapping (extractCallGraph), ensuring accurate node and edge detection.
  • Tree-sitter integration in src/plugins/tree-sitter-plugin.ts provides the initial AST that extractors traverse to identify functions, classes, imports, and calls.
  • Helper functions like isPublic, extractParams, and extractCalleeName handle language-specific semantics while shared utilities in base-extractor.ts provide common traversal logic.
  • Post-processing steps handle language quirks like Rust’s impl blocks, merging methods into their owning types before returning the final StructuralAnalysis and CallGraphEntry arrays.

Frequently Asked Questions

How does the extractor handle different programming languages uniformly?

Each extractor implements the same LanguageExtractor interface defined in src/plugins/extractors/types.ts. While the specific node types vary between languages (e.g., function_definition in Python versus function_item in Rust), the workflow remains identical: parse with tree-sitter, walk the AST to build nodes, then walk again to build edges.

What information is captured in the structural nodes?

Structural nodes contain the entity type (function, class, enum, trait), name, visibility (public/private), source line ranges, and language-specific details like parameters and return types. The StructuralAnalysis interface standardizes this data across all languages.

How are method calls resolved when the caller is not in the same file?

The extractCallGraph method captures the caller from the current function stack and the callee from the call_expression node. While the extractor records the identifier name, full cross-file resolution happens later in src/analyzer/graph-builder.ts, which merges individual file results into the unified knowledge graph.

Why does the Rust extractor use a methodsByType map?

Rust separates method definitions from struct declarations using impl blocks. The extractor uses methodsByType (lines 101-108) to temporarily store methods during the first pass, then attaches them to their owning structs or enums (lines 135-142) before returning the final structure. This ensures the knowledge graph presents methods as members of their types rather than standalone functions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →