# How Egonex-AI Parsers Handle Non-Code Files: YAML, JSON, TOML, .env, and Dockerfile Analysis

> Discover how Egonex-AI parsers analyze non-code files like YAML, JSON, TOML, env, and Dockerfile. Our AI converts configuration content into structured graph nodes for deeper understanding.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Egonex-AI treats configuration files as first-class citizens through dedicated AnalyzerPlugin implementations that convert YAML, JSON, TOML, .env, and Dockerfile content into structured graph nodes.** 

The Understand-Anything repository by Egonex-AI provides a unified parsing framework that ingests non-code files into its knowledge graph. Each format receives specialized handling via distinct parser classes located in the core plugin package, ensuring accurate structural analysis regardless of file type.

## Architecture of the AnalyzerPlugin System

Every parser in the Egonex-AI ecosystem implements the **AnalyzerPlugin** interface. This contract requires three core components: a `languages` array declaring supported file identifiers, an `analyzeFile` method receiving file paths and raw content, and optional helper methods for reference extraction.

The central registry in [`plugins/registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/plugins/registry.ts) maintains a mapping between language IDs and their corresponding plugin instances. When the core analyzer scans a project, it queries this registry using the file’s language identifier—derived from extensions or custom detection heuristics—and routes matching files to the appropriate `analyzeFile` implementation.

## YAML Parsing Strategy

The **YAMLConfigParser** class in [`understand-anything-plugin/packages/core/src/plugins/parsers/yaml-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/yaml-parser.ts) processes YAML documents using the `yaml` npm package to construct JavaScript objects. 

Top-level keys are enumerated and their line numbers located via regex pattern matching (`^["']?key["']?\s*:`). If the YAML parse fails, the parser falls back to a simpler regex (`/^(\w[\w-]*)\s*:/`) to extract keys without full parsing.

The parser also registers special flavors including **docker-compose**, **kubernetes**, **github-actions**, and **openapi** to ensure these variants are properly recognized by the language registry.

## JSON and JSONC Processing

The **JSONConfigParser** in [`understand-anything-plugin/packages/core/src/plugins/parsers/json-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/json-parser.ts) handles both standard JSON and JSON-with-Comments (JSONC) formats.

The implementation first strips comments and trailing commas using the `stripJsoncSyntax` helper before passing the cleaned string to `JSON.parse`. Top-level keys become sections with line numbers derived from searching for quoted keys in the original file. 

The parser additionally extracts `$ref` references—a pattern common in JSON Schema and OpenAPI specifications—and reports them as `ReferenceResolution` objects through the `extractReferences` method.

## TOML Section Extraction

**TOMLParser**, located in [`understand-anything-plugin/packages/core/src/plugins/parsers/toml-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/toml-parser.ts), takes a lightweight approach to TOML files.

It scans each line for section headers matching `[section]` or `[[array-of-tables]]` patterns. The header name determines nesting levels via `name.split(".").length`. Only section titles and their line ranges are emitted to the graph; key-value pairs within sections are intentionally ignored to maintain a focused structural overview.

## Environment Variable Parsing

The **EnvParser** class in [`understand-anything-plugin/packages/core/src/plugins/parsers/env-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/env-parser.ts) processes `.env` files through line-by-line iteration.

It skips comments (lines beginning with `#`) and blank lines, then matches `KEY=value` patterns using regex. Each discovered key becomes a `DefinitionInfo` entry of kind `"variable"` with its precise line range. The parser maintains a lightweight profile by not supporting `export VAR=…` syntax or multi-line values.

## Dockerfile Instruction Analysis

**DockerfileParser** in [`understand-anything-plugin/packages/core/src/plugins/parsers/dockerfile-parser.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/parsers/dockerfile-parser.ts) treats Dockerfiles as instruction sequences rather than configuration files.

The parser walks the file line-by-line, recognizing Docker instructions including `FROM`, `RUN`, `CMD`, `ENTRYPOINT`, and `ENV`. Each instruction emits as a section with its name set to the instruction keyword and a `lineRange` covering the full instruction including line-continuations (`\`). This produces a high-level skeleton of the Docker build process suitable for knowledge graph visualization.

## Usage Examples

### Direct Parser Invocation

You can instantiate parsers directly for unit testing or custom tooling:

```typescript
import { YAMLConfigParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

async function dumpYamlSections(path: string) {
  const content = await readFile(path, 'utf-8');
  const parser = new YAMLConfigParser();
  const analysis = parser.analyzeFile(path, content);
  console.log('YAML sections:', analysis.sections);
}

```

### Full Project Analysis

The CLI automatically handles mixed projects through the core analyzer:

```typescript
import { analyzeProject } from '@understand-anything/core';
import { resolve } from 'path';

const projectRoot = resolve(process.cwd(), 'my-project');
await analyzeProject(projectRoot);   // discovers .yaml, .jsonc, .toml, .env, Dockerfile, etc.

```

### Extracting JSON Schema References

For JSON Schema files containing external references:

```typescript
import { JSONConfigParser } from '@understand-anything/core';
import { readFile } from 'fs/promises';

const content = await readFile('schema.json', 'utf-8');
const parser = new JSONConfigParser();
const refs = parser.extractReferences('schema.json', content);
console.log('External refs:', refs);

```

## Summary

- **Egonex-AI parsers** treat YAML, JSON, TOML, .env, and Dockerfile formats as analyzable graph nodes through the `AnalyzerPlugin` interface.
- Each format has a dedicated parser class in `understand-anything-plugin/packages/core/src/plugins/parsers/` that implements the `analyzeFile` method.
- The **registry** in [`plugins/registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/plugins/registry.ts) routes files to appropriate parsers based on language ID detection.
- **YAML parser** handles special flavors like Kubernetes and Docker Compose while providing regex fallbacks for malformed files.
- **JSON parser** strips JSONC syntax and extracts `$ref` references for schema linking.
- **TOML parser** focuses exclusively on section headers to define structural hierarchy.
- **Env and Dockerfile parsers** use line-based scanning to extract variables and build instructions respectively.

## Frequently Asked Questions

### How does Egonex-AI handle malformed YAML files?

If the `yaml` npm package fails to parse a file, the `YAMLConfigParser` falls back to a simple regex pattern (`/^(\w[\w-]*)\s*:/`) that extracts top-level keys without requiring valid YAML syntax. This ensures partial analysis even when configuration files contain syntax errors.

### Can Egonex-AI parsers extract references between JSON Schema files?

Yes. The `JSONConfigParser` specifically identifies `$ref` pointers within JSON content and returns them as `ReferenceResolution` objects through the `extractReferences` method. This enables cross-file navigation in OpenAPI and JSON Schema documents.

### Does the TOML parser extract individual key-value pairs?

No. The `TOMLParser` intentionally limits extraction to section headers (`[section]` and `[[array-of-tables]]`) and their line ranges. It calculates nesting depth from dot-separated header names but does not emit individual key-value definitions to the knowledge graph.

### What Dockerfile instructions does the parser recognize?

The `DockerfileParser` recognizes standard Docker instructions including `FROM`, `RUN`, `CMD`, `ENTRYPOINT`, `ENV`, and others. Each instruction becomes a graph section with its line range calculated to include line continuations (`\`), providing a complete skeleton of the container build process.