# How the Mermaid Import Parser Handles Edge Cases in Diagram Generation

> Discover how the Mermaid import parser handles edge cases in diagram generation. Learn about its ten-phase pipeline, strict boundaries, and efficient resource management.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: internals
- Published: 2026-09-11

---

**The Mermaid import parser in `cathrynlavery/diagram-design` employs a ten-phase defensive pipeline with strict trust boundaries, failing fast on malformed input while bounding resource usage to 4 MiB and preserving only semantic diagram structure.**

The `cathrynlavery/diagram-design` repository contains a robust import parser located at [`skills/diagram-design/scripts/mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/mermaid_extract.py) that transforms Mermaid source files into a safe intermediate representation. This **Mermaid import parser** guarantees semantic-only extraction through a series of defensive parsing stages designed to handle oversized files, unterminated fences, unsupported grammars, and malformed edges without compromising system stability.

## Input Normalization and Resource Bounding

The parser establishes a strict trust boundary during **Phase 1: Input Normalisation**. It reads source files (plain `.mmd`, `.mermaid`, or fenced Markdown) with a hard size limit of `MAX_SOURCE_BYTES = 4 MiB` and decodes content as UTF-8.

If a source file exceeds this limit, the parser immediately triggers `node limit exceeded` or `source exceeds … MiB` errors (see lines 31‑34 of [`mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/mermaid_extract.py)). This prevents denial-of-service attacks via massive file uploads and ensures predictable memory usage.

## Block Detection and Front-Matter Isolation

**Phase 2: Block Detection** uses the `load_blocks()` function to extract either entire files (for `.mmd`/`.mermaid` extensions) or individual fenced Mermaid blocks from Markdown documents. The parser exits with code 2 if no fenced block is found or if it encounters an unterminated fence (lines 56‑60).

Edge cases handled here include:
- **Missing fences**: Returns `no fenced mermaid block found`
- **Unterminated fences**: Detected during block scanning

**Phase 3: Front-Matter Skipping** occurs in `_prepared_lines()`, which discards any leading YAML-style front-matter (`--- … ---`) before grammar detection (lines 88‑94). This ensures that front-matter metadata never leaks into diagram kind or direction detection, verified in [`verify_mermaid_import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/verify_mermaid_import.py) lines 64‑80.

## Grammar Validation and Statement Assembly

During **Phase 4: Grammar & Direction Detection**, the `_kind_and_direction()` function recognizes four supported grammars: `flowchart`/`graph`, `sequenceDiagram`, `stateDiagram-v2`, and `erDiagram`. Unsupported kinds (such as `pie` or `mindmap`) cause immediate failure with a clear `unsupported diagram kind` message (lines 28‑33, verified at lines 64‑70).

**Phase 5: Tokenisation & Statement Assembly** handles multiline complexity through `_logical_statements()`. This function:
- Joins multiline statements into logical units
- Validates balanced quotes and brackets via `_statement_complete()`
- Splits statements on top-level semicolons

When encountering unterminated statements, the parser raises `malformed edge at line …` and aborts (lines 29‑31), preventing partially parsed corrupt data from entering the pipeline.

## Robust Node and Edge Parsing

**Phase 6: Node Parsing** via `_parse_node_expression()` strips class suffixes (e.g., `:::warning`), expands `@{ … }` attribute blocks, classifies shapes through `classify_shape()`, and normalizes labels using `clean_label()`. Class suffixes are verified as stripped in [`verify_mermaid_import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/verify_mermaid_import.py) lines 66‑70, and expanded attributes are normalized to known families (lines 26‑55).

**Phase 7: Edge Parsing** employs `_edge_operators()` to build a masking system that blanks quoted text and brackets before matching. This architecture handles:
- **Spaced edges**: `A-- “text” -->B`
- **Compact edges**: `A--yes-->B`
- **Bidirectional arrows**: `<<-->>`, `o--`, `x--`
- **Undirected markers** and operator styles (solid, dashed, thick)

Labels are extracted from raw source between operators, with all variations exercised in the verification suite sections "compact edge labels" and "quoted participants".

## Semantic Purification and Cycle Detection

**Phase 8: Semantic Validation** removes non-semantic directives through `_discard_nonsemantic()`. This strips `style`, `classDef`, and `click` handlers while tallying discarded items in `diagram.discarded` for the fidelity ledger (verification lines 96‑104).

**Phase 9: Graph Analysis** computes structural properties via `_finalize_degrees()` (calculating in/out-degrees) and `_has_cycle()` (running DFS cycle detection). This guarantees detection of circular references such as self-loops in flowcharts (verified at lines 80‑86).

## Safe Output Generation

**Phase 10: Output Generation** provides two export paths:
- `digest()`: Creates human-readable Markdown reports
- `to_json()`: Emits the full intermediate representation

Both respect the `--max-rows` limit and never re-encode labels (preserving plain text). Budget checks for `over_node_budget` and `over_edge_budget` appear in the digest (lines 71‑79), ensuring output never exceeds configured display limits.

## Working with the Parser

Extract a diagram to JSON using the command line:

```bash
python3 skills/diagram-design/scripts/mermaid_extract.py docs/fixtures/sample-flowchart.mmd --json

```

Import the parser as a module for programmatic access:

```python
from importlib.util import spec_from_file_location, module_from_spec

spec = spec_from_file_location("mermaid_extract", "skills/diagram-design/scripts/mermaid_extract.py")
mermaid = module_from_spec(spec)
spec.loader.exec_module(mermaid)

# Run the extractor on a file and capture the digest

result = mermaid.main(["scripts/fixtures/sample-adversarial.mmd"])

```

The parser preserves UTF-8 content such as multilingual labels:

```mermaid
flowchart LR
    A["登录<br/>続行 ⇒"] --> B["résumé"]

```

Running the extractor on this snippet demonstrates UTF-8 preservation as verified in [`verify_mermaid_import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/verify_mermaid_import.py) lines 68‑95.

## Core Source Files

| File | Role |
|------|------|
| [`skills/diagram-design/scripts/mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/mermaid_extract.py) | Core parser that normalises Mermaid source into a safe IR |
| [`scripts/verify-mermaid-import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-mermaid-import.py) | Comprehensive test harness validating edge-case handling |
| [`skills/diagram-design/references/import-mermaid.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/import-mermaid.md) | User-facing documentation for import workflows |
| [`skills/diagram-design/SKILL.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/SKILL.md) | Skill capability declarations and extractor wiring |
| `scripts/fixtures/sample-adversarial.mmd` | Challenging test inputs for the verification suite |

## Summary

- The **Mermaid import parser** implements a ten-phase defensive pipeline with strict size limits (`MAX_SOURCE_BYTES = 4 MiB`) and UTF-8 preservation.
- It **fails fast** on unsupported grammars, unterminated fences, and malformed edges while providing clear error messages and specific exit codes.
- **Semantic purification** removes non-structural directives (`style`, `classDef`, `click`) while preserving diagram topology.
- **Cycle detection** via DFS ensures structural integrity before output generation.
- Output is bounded by `--max-rows` and includes fidelity ledgers showing discarded content.

## Frequently Asked Questions

### What happens when the Mermaid import parser encounters an unsupported diagram type?

The parser immediately aborts during Phase 4 with an `unsupported diagram kind` message. According to lines 28‑33 of [`mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/mermaid_extract.py), only `flowchart`, `sequenceDiagram`, `stateDiagram-v2`, and `erDiagram` are recognized; all other grammars trigger an early exit before tokenization begins.

### How does the parser handle Mermaid files with YAML front-matter?

The `_prepared_lines()` function in Phase 3 scans for leading `---` delimiters and discards everything up to the closing `---` before grammar detection (lines 88‑94). This prevents front-matter metadata from interfering with diagram type recognition or node parsing, as verified in [`verify_mermaid_import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/verify_mermaid_import.py) lines 64‑80.

### Can the parser handle edge labels containing quotes or Unicode characters?

Yes. The `_edge_operators()` function builds a masking system that blanks quoted text before parsing operators, supporting both spaced (`A-- "text" -->B`) and compact (`A--label-->B`) syntax. The parser decodes all input as UTF-8 and never re-encodes labels during output generation, preserving multilingual characters and symbols as demonstrated in the verification suite (lines 68‑95).

### What resource limits protect the system from oversized diagram files?

The parser enforces a hard limit of `MAX_SOURCE_BYTES = 4 MiB` during Phase 1. Files exceeding this threshold trigger `source exceeds … MiB` or `node limit exceeded` errors (lines 31‑34), preventing memory exhaustion attacks and ensuring the extractor runs in bounded space.