How the Mermaid Import Parser Handles Edge Cases in Diagram Generation
The Mermaid import parser in cathrynlavery/diagram-design employs a ten-phase defensive pipeline with strict trust boundaries, failing fast on malformed input while bounding resource usage to 4 MiB and preserving only semantic diagram structure.
The cathrynlavery/diagram-design repository contains a robust import parser located at skills/diagram-design/scripts/mermaid_extract.py that transforms Mermaid source files into a safe intermediate representation. This Mermaid import parser guarantees semantic-only extraction through a series of defensive parsing stages designed to handle oversized files, unterminated fences, unsupported grammars, and malformed edges without compromising system stability.
Input Normalization and Resource Bounding
The parser establishes a strict trust boundary during Phase 1: Input Normalisation. It reads source files (plain .mmd, .mermaid, or fenced Markdown) with a hard size limit of MAX_SOURCE_BYTES = 4 MiB and decodes content as UTF-8.
If a source file exceeds this limit, the parser immediately triggers node limit exceeded or source exceeds … MiB errors (see lines 31‑34 of mermaid_extract.py). This prevents denial-of-service attacks via massive file uploads and ensures predictable memory usage.
Block Detection and Front-Matter Isolation
Phase 2: Block Detection uses the load_blocks() function to extract either entire files (for .mmd/.mermaid extensions) or individual fenced Mermaid blocks from Markdown documents. The parser exits with code 2 if no fenced block is found or if it encounters an unterminated fence (lines 56‑60).
Edge cases handled here include:
- Missing fences: Returns
no fenced mermaid block found - Unterminated fences: Detected during block scanning
Phase 3: Front-Matter Skipping occurs in _prepared_lines(), which discards any leading YAML-style front-matter (--- … ---) before grammar detection (lines 88‑94). This ensures that front-matter metadata never leaks into diagram kind or direction detection, verified in verify_mermaid_import.py lines 64‑80.
Grammar Validation and Statement Assembly
During Phase 4: Grammar & Direction Detection, the _kind_and_direction() function recognizes four supported grammars: flowchart/graph, sequenceDiagram, stateDiagram-v2, and erDiagram. Unsupported kinds (such as pie or mindmap) cause immediate failure with a clear unsupported diagram kind message (lines 28‑33, verified at lines 64‑70).
Phase 5: Tokenisation & Statement Assembly handles multiline complexity through _logical_statements(). This function:
- Joins multiline statements into logical units
- Validates balanced quotes and brackets via
_statement_complete() - Splits statements on top-level semicolons
When encountering unterminated statements, the parser raises malformed edge at line … and aborts (lines 29‑31), preventing partially parsed corrupt data from entering the pipeline.
Robust Node and Edge Parsing
Phase 6: Node Parsing via _parse_node_expression() strips class suffixes (e.g., :::warning), expands @{ … } attribute blocks, classifies shapes through classify_shape(), and normalizes labels using clean_label(). Class suffixes are verified as stripped in verify_mermaid_import.py lines 66‑70, and expanded attributes are normalized to known families (lines 26‑55).
Phase 7: Edge Parsing employs _edge_operators() to build a masking system that blanks quoted text and brackets before matching. This architecture handles:
- Spaced edges:
A-- “text” -->B - Compact edges:
A--yes-->B - Bidirectional arrows:
<<-->>,o--,x-- - Undirected markers and operator styles (solid, dashed, thick)
Labels are extracted from raw source between operators, with all variations exercised in the verification suite sections "compact edge labels" and "quoted participants".
Semantic Purification and Cycle Detection
Phase 8: Semantic Validation removes non-semantic directives through _discard_nonsemantic(). This strips style, classDef, and click handlers while tallying discarded items in diagram.discarded for the fidelity ledger (verification lines 96‑104).
Phase 9: Graph Analysis computes structural properties via _finalize_degrees() (calculating in/out-degrees) and _has_cycle() (running DFS cycle detection). This guarantees detection of circular references such as self-loops in flowcharts (verified at lines 80‑86).
Safe Output Generation
Phase 10: Output Generation provides two export paths:
digest(): Creates human-readable Markdown reportsto_json(): Emits the full intermediate representation
Both respect the --max-rows limit and never re-encode labels (preserving plain text). Budget checks for over_node_budget and over_edge_budget appear in the digest (lines 71‑79), ensuring output never exceeds configured display limits.
Working with the Parser
Extract a diagram to JSON using the command line:
python3 skills/diagram-design/scripts/mermaid_extract.py docs/fixtures/sample-flowchart.mmd --json
Import the parser as a module for programmatic access:
from importlib.util import spec_from_file_location, module_from_spec
spec = spec_from_file_location("mermaid_extract", "skills/diagram-design/scripts/mermaid_extract.py")
mermaid = module_from_spec(spec)
spec.loader.exec_module(mermaid)
# Run the extractor on a file and capture the digest
result = mermaid.main(["scripts/fixtures/sample-adversarial.mmd"])
The parser preserves UTF-8 content such as multilingual labels:
flowchart LR
A["登录<br/>続行 ⇒"] --> B["résumé"]
Running the extractor on this snippet demonstrates UTF-8 preservation as verified in verify_mermaid_import.py lines 68‑95.
Core Source Files
| File | Role |
|---|---|
skills/diagram-design/scripts/mermaid_extract.py |
Core parser that normalises Mermaid source into a safe IR |
scripts/verify-mermaid-import.py |
Comprehensive test harness validating edge-case handling |
skills/diagram-design/references/import-mermaid.md |
User-facing documentation for import workflows |
skills/diagram-design/SKILL.md |
Skill capability declarations and extractor wiring |
scripts/fixtures/sample-adversarial.mmd |
Challenging test inputs for the verification suite |
Summary
- The Mermaid import parser implements a ten-phase defensive pipeline with strict size limits (
MAX_SOURCE_BYTES = 4 MiB) and UTF-8 preservation. - It fails fast on unsupported grammars, unterminated fences, and malformed edges while providing clear error messages and specific exit codes.
- Semantic purification removes non-structural directives (
style,classDef,click) while preserving diagram topology. - Cycle detection via DFS ensures structural integrity before output generation.
- Output is bounded by
--max-rowsand includes fidelity ledgers showing discarded content.
Frequently Asked Questions
What happens when the Mermaid import parser encounters an unsupported diagram type?
The parser immediately aborts during Phase 4 with an unsupported diagram kind message. According to lines 28‑33 of mermaid_extract.py, only flowchart, sequenceDiagram, stateDiagram-v2, and erDiagram are recognized; all other grammars trigger an early exit before tokenization begins.
How does the parser handle Mermaid files with YAML front-matter?
The _prepared_lines() function in Phase 3 scans for leading --- delimiters and discards everything up to the closing --- before grammar detection (lines 88‑94). This prevents front-matter metadata from interfering with diagram type recognition or node parsing, as verified in verify_mermaid_import.py lines 64‑80.
Can the parser handle edge labels containing quotes or Unicode characters?
Yes. The _edge_operators() function builds a masking system that blanks quoted text before parsing operators, supporting both spaced (A-- "text" -->B) and compact (A--label-->B) syntax. The parser decodes all input as UTF-8 and never re-encodes labels during output generation, preserving multilingual characters and symbols as demonstrated in the verification suite (lines 68‑95).
What resource limits protect the system from oversized diagram files?
The parser enforces a hard limit of MAX_SOURCE_BYTES = 4 MiB during Phase 1. Files exceeding this threshold trigger source exceeds … MiB or node limit exceeded errors (lines 31‑34), preventing memory exhaustion attacks and ensuring the extractor runs in bounded space.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →