How JSON, YAML, Markdown, Docker, and SQL Parser Plugins Integrate with Tree-Sitter

The JSON, YAML, Markdown, Docker, and SQL parser plugins in Understand-Anything bridge lightweight custom parsers with the core Tree-Sitter pipeline by implementing a unified TreeSitterParser interface that produces AST-compatible node trees, enabling uniform structural analysis across both grammar-rich languages and data-format files.

The Understand-Anything repository leverages Tree-Sitter via the web-tree-sitter WASM package to generate abstract syntax trees for code analysis. While languages like TypeScript or Python utilize full Tree-Sitter grammars, data formats including JSON, YAML, Markdown, Dockerfiles, and SQL rely on specialized parser plugins that integrate seamlessly with the same extraction pipeline located in packages/core/src/plugins/.

The Tree-Sitter Foundation

At the architecture's core, the TreeSitterPlugin class in packages/core/src/plugins/tree-sitter-plugin.ts manages all language parsing through WebAssembly. It performs dynamic imports of web-tree-sitter, reads the wasmPackage and wasmFile properties from language configuration files such as json-config.ts, and instantiates a Parser object for each supported language.

When a file enters the analysis pipeline, the plugin loads the appropriate WASM grammar to create a standard Tree-Sitter AST. However, for the five data-format languages, the system employs a hybrid approach where the plugin loader still coordinates parsing, but the actual node tree construction delegates to lightweight custom implementations that mimic the WASM grammar output structure.

Lightweight Parser Plugin Architecture

The five specialized plugins reside in packages/core/src/plugins/parsers/ and handle languages lacking dedicated Tree-Sitter grammars. Each plugin exports a parse function conforming to the TreeSitterParser interface expected by the core's extractor system in src/plugins/extractors/*-extractor.ts.

Rather than compiling WASM binaries, these plugins implement hand-rolled tokenizers that produce node structures matching the shape of true Tree-Sitter trees. This design allows downstream analysis—such as dependency extraction and call-graph building—to process all file types uniformly, regardless of whether the AST originated from a compiled grammar or a custom tokenizer.

JSON Parser Implementation

The json-parser.ts file wraps JSON.parse in a Tree-Sitter-compatible node structure. It transforms the parsed JavaScript object into a hierarchical tree that the extractor pipeline traverses identically to a compiled grammar AST, capturing key-value structures and nested objects.

YAML Parser Implementation

Similarly, yaml-parser.ts implements a simplified YAML tokenizer that converts document structures into normalized nodes. This approach handles YAML's indentation-based syntax without requiring the overhead of a full Tree-Sitter grammar binary.

Markdown Parser Implementation

The markdown-parser.ts plugin breaks documents into block-level nodes—including headings, code fences, and paragraph blocks—creating a structural tree suitable for documentation analysis and cross-referencing within the knowledge graph.

Dockerfile Parser Implementation

For Dockerfiles, dockerfile-parser.ts implements a directive-aware line parser that identifies instructions such as FROM, RUN, and COPY. It constructs a minimal AST capturing image dependencies and build steps, feeding this structure into the standard extraction pipeline.

SQL Parser Implementation

The sql-parser.ts plugin employs a simple statement splitter to divide SQL scripts into individual query nodes. This enables the system to analyze database schemas and query dependencies without compiling a heavy SQL grammar into WASM.

The Three-Step Integration Process

The integration between lightweight plugins and the Tree-Sitter core follows a standardized three-phase pipeline:

  1. Language Loading: TreeSitterPlugin resolves the language identifier through configuration files (e.g., yaml-config.ts), initializing the parsing environment and loading WASM when available.

  2. Parser Registration: Each plugin registers its parse function, which accepts source text and an optional Tree-Sitter context, returning a standardized node tree. The core's extractor code receives this output regardless of whether it came from web-tree-sitter or a custom tokenizer.

  3. Uniform Extraction: The resulting AST feeds into the same structural analysis pipeline. The graph builder processes functions, classes, imports, or data structure definitions uniformly across all file types, treating lightweight parser output identically to WASM-generated trees.

Practical Implementation Examples

The following examples demonstrate how these plugins operate within the Understand-Anything framework.

For JSON files, the integration leverages the Tree-Sitter loader while substituting the custom parser:

import { TreeSitterPlugin } from "@understand-anything/core/plugins/tree-sitter-plugin";
import { jsonParser } from "@understand-anything/core/plugins/parsers/json-parser";

async function analyseJson(source: string) {
  // TreeSitterPlugin loads the generic TS/JS grammar (used for structural analysis)
  const treeSitter = await TreeSitterPlugin.loadLanguage("json"); // loads WASM if present
  // jsonParser.parse returns a Tree‑Sitter‑compatible node tree
  const ast = jsonParser.parse(source, treeSitter);
  // The extractor can now walk `ast` just like any other language
  return extractStructuralInfo(ast);
}

For Dockerfile analysis, where no WASM grammar exists, the plugin operates standalone:

import { dockerfileParser } from "@understand-anything/core/plugins/parsers/dockerfile-parser";

function analyseDockerfile(src: string) {
  // dockerfileParser builds a small node tree from Dockerfile directives
  const ast = dockerfileParser.parse(src);
  // Hand‑off to the generic extractor
  return extractStructuralInfo(ast);
}

Summary

  • The TreeSitterPlugin in packages/core/src/plugins/tree-sitter-plugin.ts coordinates all parsing through the web-tree-sitter WASM infrastructure.
  • JSON, YAML, Markdown, Dockerfile, and SQL parsers reside in packages/core/src/plugins/parsers/ and implement the TreeSitterParser interface.
  • These plugins produce AST-compatible node trees using lightweight tokenizers rather than full grammars, enabling processing by the unified extractor pipeline in src/plugins/extractors/.
  • Language configurations in files like json-config.ts declare WASM dependencies, while parser plugins handle actual node construction for data-format files.
  • The architecture ensures uniform structural analysis across both grammar-rich languages and configuration files.

Frequently Asked Questions

Why doesn't Understand-Anything use standard Tree-Sitter grammars for JSON and YAML?

While Tree-Sitter grammars exist for these languages, the project employs lightweight custom parsers to minimize WASM binary size and startup overhead. The json-parser.ts and yaml-parser.ts implementations provide sufficient structural information for dependency analysis and graph building without the complexity of maintaining grammar compilations for simple data formats.

How does the extractor pipeline handle ASTs from different parser sources?

The extractors in src/plugins/extractors/*-extractor.ts operate on a normalized node interface defined by the TreeSitterParser contract. Whether the AST originates from web-tree-sitter WASM or a custom tokenizer like dockerfile-parser.ts, the nodes expose consistent properties for traversal, allowing the same extraction logic to process imports, dependencies, and structural elements uniformly.

Can custom parser plugins support languages beyond the five mentioned?

Yes. The architecture supports extending the parser directory with new implementations. Any plugin exporting a parse function that returns Tree-Sitter-compatible nodes can integrate with the TreeSitterPlugin loader and participate in the extraction pipeline, provided it registers a language identifier in the configuration system using wasmPackage declarations.

What performance implications exist for the hybrid parsing approach?

The lightweight parsers for JSON (JSON.parse), YAML, and Dockerfiles execute significantly faster than WASM loading for simple files, reducing cold-start latency. However, complex SQL parsing in sql-parser.ts may incur overhead compared to compiled grammars for large schemas, though the unified interface ensures this remains transparent to downstream analysis components.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →