How Egonex-AI Parsers Handle Non-Code Files: YAML, JSON, TOML, .env, and Dockerfile Analysis

Egonex-AI treats configuration files as first-class citizens through dedicated AnalyzerPlugin implementations that convert YAML, JSON, TOML, .env, and Dockerfile content into structured graph nodes.

The Understand-Anything repository by Egonex-AI provides a unified parsing framework that ingests non-code files into its knowledge graph. Each format receives specialized handling via distinct parser classes located in the core plugin package, ensuring accurate structural analysis regardless of file type.

Architecture of the AnalyzerPlugin System

Every parser in the Egonex-AI ecosystem implements the AnalyzerPlugin interface. This contract requires three core components: a languages array declaring supported file identifiers, an analyzeFile method receiving file paths and raw content, and optional helper methods for reference extraction.

The central registry in plugins/registry.ts maintains a mapping between language IDs and their corresponding plugin instances. When the core analyzer scans a project, it queries this registry using the file’s language identifier—derived from extensions or custom detection heuristics—and routes matching files to the appropriate analyzeFile implementation.

YAML Parsing Strategy

The YAMLConfigParser class in understand-anything-plugin/packages/core/src/plugins/parsers/yaml-parser.ts processes YAML documents using the yaml npm package to construct JavaScript objects.

Top-level keys are enumerated and their line numbers located via regex pattern matching (^["']?key["']?\s*:). If the YAML parse fails, the parser falls back to a simpler regex (/^(\w[\w-]*)\s*:/) to extract keys without full parsing.

The parser also registers special flavors including docker-compose, kubernetes, github-actions, and openapi to ensure these variants are properly recognized by the language registry.

JSON and JSONC Processing

The JSONConfigParser in understand-anything-plugin/packages/core/src/plugins/parsers/json-parser.ts handles both standard JSON and JSON-with-Comments (JSONC) formats.

The implementation first strips comments and trailing commas using the stripJsoncSyntax helper before passing the cleaned string to JSON.parse. Top-level keys become sections with line numbers derived from searching for quoted keys in the original file.

The parser additionally extracts $ref references—a pattern common in JSON Schema and OpenAPI specifications—and reports them as ReferenceResolution objects through the extractReferences method.

TOML Section Extraction

TOMLParser, located in understand-anything-plugin/packages/core/src/plugins/parsers/toml-parser.ts, takes a lightweight approach to TOML files.

It scans each line for section headers matching [section] or [[array-of-tables]] patterns. The header name determines nesting levels via name.split(".").length. Only section titles and their line ranges are emitted to the graph; key-value pairs within sections are intentionally ignored to maintain a focused structural overview.

Environment Variable Parsing

The EnvParser class in understand-anything-plugin/packages/core/src/plugins/parsers/env-parser.ts processes .env files through line-by-line iteration.

It skips comments (lines beginning with #) and blank lines, then matches KEY=value patterns using regex. Each discovered key becomes a DefinitionInfo entry of kind "variable" with its precise line range. The parser maintains a lightweight profile by not supporting export VAR=… syntax or multi-line values.

Dockerfile Instruction Analysis

DockerfileParser in understand-anything-plugin/packages/core/src/plugins/parsers/dockerfile-parser.ts treats Dockerfiles as instruction sequences rather than configuration files.

The parser walks the file line-by-line, recognizing Docker instructions including FROM, RUN, CMD, ENTRYPOINT, and ENV. Each instruction emits as a section with its name set to the instruction keyword and a lineRange covering the full instruction including line-continuations (\). This produces a high-level skeleton of the Docker build process suitable for knowledge graph visualization.

Usage Examples

Direct Parser Invocation

You can instantiate parsers directly for unit testing or custom tooling:

import { YAMLConfigParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

async function dumpYamlSections(path: string) {
  const content = await readFile(path, 'utf-8');
  const parser = new YAMLConfigParser();
  const analysis = parser.analyzeFile(path, content);
  console.log('YAML sections:', analysis.sections);
}

Full Project Analysis

The CLI automatically handles mixed projects through the core analyzer:

import { analyzeProject } from '@understand-anything/core';
import { resolve } from 'path';

const projectRoot = resolve(process.cwd(), 'my-project');
await analyzeProject(projectRoot);   // discovers .yaml, .jsonc, .toml, .env, Dockerfile, etc.

Extracting JSON Schema References

For JSON Schema files containing external references:

import { JSONConfigParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

const content = await readFile('schema.json', 'utf-8');
const parser = new JSONConfigParser();
const refs = parser.extractReferences('schema.json', content);
console.log('External refs:', refs);

Summary

  • Egonex-AI parsers treat YAML, JSON, TOML, .env, and Dockerfile formats as analyzable graph nodes through the AnalyzerPlugin interface.
  • Each format has a dedicated parser class in understand-anything-plugin/packages/core/src/plugins/parsers/ that implements the analyzeFile method.
  • The registry in plugins/registry.ts routes files to appropriate parsers based on language ID detection.
  • YAML parser handles special flavors like Kubernetes and Docker Compose while providing regex fallbacks for malformed files.
  • JSON parser strips JSONC syntax and extracts $ref references for schema linking.
  • TOML parser focuses exclusively on section headers to define structural hierarchy.
  • Env and Dockerfile parsers use line-based scanning to extract variables and build instructions respectively.

Frequently Asked Questions

How does Egonex-AI handle malformed YAML files?

If the yaml npm package fails to parse a file, the YAMLConfigParser falls back to a simple regex pattern (/^(\w[\w-]*)\s*:/) that extracts top-level keys without requiring valid YAML syntax. This ensures partial analysis even when configuration files contain syntax errors.

Can Egonex-AI parsers extract references between JSON Schema files?

Yes. The JSONConfigParser specifically identifies $ref pointers within JSON content and returns them as ReferenceResolution objects through the extractReferences method. This enables cross-file navigation in OpenAPI and JSON Schema documents.

Does the TOML parser extract individual key-value pairs?

No. The TOMLParser intentionally limits extraction to section headers ([section] and [[array-of-tables]]) and their line ranges. It calculates nesting depth from dot-separated header names but does not emit individual key-value definitions to the knowledge graph.

What Dockerfile instructions does the parser recognize?

The DockerfileParser recognizes standard Docker instructions including FROM, RUN, CMD, ENTRYPOINT, ENV, and others. Each instruction becomes a graph section with its line range calculated to include line continuations (\), providing a complete skeleton of the container build process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →