# Mermaid Import Pipeline for Multi-Block Markdown and Adversarial Labels: A Complete Technical Guide

> Explore the secure three-stage Mermaid import pipeline from cathrynlavery/diagram-design. Extract, sanitize, and redraw diagrams from Markdown with adversarial label normalization. Get the complete technical guide.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The Diagram-Design repository implements a secure three-stage pipeline that extracts, sanitizes, and redraws Mermaid diagrams from Markdown files while treating all input as untrusted data and normalizing adversarial labels to plain text.**

The **Mermaid import pipeline** enables the Diagram-Design skill to transform existing Mermaid sources—whether standalone `.mmd` files or fenced blocks within Markdown—into editorial-quality diagrams. According to the cathrynlavery/diagram-design source code, this pipeline enforces a strict trust boundary that never executes, renders, or fetches external content, ensuring safe processing of potentially hostile inputs.

## The Three-Stage Import Architecture

The pipeline processes all Mermaid sources through three logical stages defined in [`references/import-mermaid.md`](https://github.com/cathrynlavery/diagram-design/blob/main/references/import-mermaid.md). Each stage operates sequentially to transform raw Mermaid syntax into a self-contained, redacted intermediate representation (IR) before final rendering.

### Stage 1: Extracting the Intermediate Representation

[`scripts/mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/mermaid_extract.py) serves as the core IR extractor, implementing the trust boundary between untrusted input and safe internal data structures. The script accepts `.mmd`, `.mermaid`, or any Markdown suffix and performs the following operations:

- **Block extraction**: For Markdown files, it locates fenced Mermaid blocks (`` ```mermaid ``) and raises an `unterminated mermaid fence` error for unterminated fences.
- **Pre-processing**: Strips front-matter, Mermaid directives (`%%{…}%%`), and styling commands (`style`, `classDef`, `class`, `linkStyle`, `click`), counting discarded elements in `diagram.discarded`.
- **Adversarial label handling**: The `clean_label` function normalizes quoted strings, HTML entities, and escaped characters—converting labels like `"yes, && drop <script>alert(1)</script>"` into safe plain text (`yes, && drop alert(1)`).
- **Grammar detection**: Identifies diagram kinds (`flowchart`, `sequenceDiagram`, `stateDiagram-v2`, `erDiagram`) and directions (`TD`, `LR`), rejecting unsupported types (pie, mindmap) with clear errors.
- **Budget analysis**: The `analyze()` method computes node/edge counts and flags diagrams exceeding the editorial budget (>9 nodes or >12 edges).

### Stage 2: Configuring the Four Dials

The import command (`/diagram-design:import-mermaid`) forwards user arguments through `prompts/import-mermaid.md` to establish output parameters defined in `references/output-spec.md`:

- **Format**: `html`, `svg`, `png`, or `html+png`
- **Size**: Presets including `doc-inline`, `slide-16x9`, `social-og`, `print-a4-landscape`, or `fit`
- **Detail**: `faithful`, `balanced`, or `simplified` (validated against the IR's `budget:` line)
- **Audience**: `engineer`, `mixed`, or `executive` (guides label phrasing)
- **Variant**: Optional `light`, `dark`, or `full` color schemes

Missing dials default to `html` and `doc-inline` unless material ambiguity requires user clarification.

### Stage 3: Redrawing with Editorial Layout

With a validated IR and configured dials, the skill generates fresh layouts starting from a blank `viewBox`—explicitly ignoring Mermaid-provided coordinates. Key transformations include:

- Converting flowchart rhombuses to decision diamonds and cylinders to *Store/State* boxes
- Mapping subgraphs to containers or collapsible groups for "quiet frames"
- Applying connector rules from `SKILL.md §6` while discarding length markers and style hints
- Generating a **fidelity ledger** documenting every merge, collapse, or drop (e.g., "9 IR nodes → 6 drawn, 1 cycle collapsed")

## Processing Multi-Block Markdown Files

The pipeline handles complex documentation containing multiple diagram definitions through indexed extraction. When processing a Markdown file containing several fenced Mermaid blocks, `mermaid_extract.py` treats each block as a separate diagram entity.

To extract all diagrams from a multi-block file:

```bash
python3 skills/diagram-design/scripts/mermaid_extract.py docs/complex-diagrams.md --diagram all --max-rows 20

```

This command generates independent digests for each block, prefixed with indices (`complex-diagrams-0`, `complex-diagrams-1`, etc.), enabling batch transformation of entire documentation suites.

When invoking the skill directly on multi-block sources:

```bash
/diagram-design:import-mermaid docs/complex-diagrams.md --diagram all --format html

```

The system produces numbered output files (`complex-diagrams-0.html`, `complex-diagrams-1.html`) preserving the original sequence while applying consistent styling across all diagrams.

## Security: Defending Against Adversarial Labels

The adversarial label handling in `scripts/mermaid_extract.py` implements defense-in-depth for untrusted Mermaid sources. The parser specifically targets attack vectors including:

- **HTML injection**: Strips `<script>` tags and entity-encodes reserved characters
- **Quote escaping**: Normalizes escaped quotes within quoted edge labels
- **Whitespace attacks**: Collapses malicious whitespace padding
- **Class suffix poisoning**: Removes `:::class` syntax that could trigger unexpected styling
- **New-style attributes**: Expands and sanitizes `@{ … }` node attribute blocks

For example, an edge definition containing:

```mermaid
A-- "yes, && drop <script>alert(1)</script>" -->B

```

Produces a sanitized IR label of `yes, && drop alert(1)` with all executable content removed.

The test harness `scripts/verify-mermaid-import.py` validates this behavior against `scripts/fixtures/sample-adversarial.mmd`, ensuring the extractor correctly sanitizes deliberately hostile constructs before they reach the redraw stage.

## Command-Line Usage Examples

### Extract JSON IR for Programmatic Consumption

```bash
python3 skills/diagram-design/scripts/mermaid_extract.py designs/architecture.mmd --json

```

Outputs a machine-readable IR including node tables, edge relationships, and budget compliance flags.

### Import with Executive Formatting

```bash
/diagram-design:import-mermaid designs/architecture.mmd --size slide-16x9 --detail balanced --audience executive

```

This executes the full pipeline: extraction, dial validation, and generation of `architecture-0.html` optimized for executive presentation.

### Handling Standalone Mermaid Files

```bash
/diagram-design:import-mermaid scripts/fixtures/sample-flowchart.mmd --format svg --size doc-wide

```

Converts the sample fixture directly to scalable vector graphics suitable for documentation embedding.

## Key Implementation Files

| File | Role |
|------|------|
| `skills/diagram-design/scripts/mermaid_extract.py` | Core IR extractor; parses Mermaid syntax, enforces trust boundary, handles adversarial labels |
| `skills/diagram-design/prompts/import-mermaid.md` | User-facing command syntax and dial definitions |
| `skills/diagram-design/references/import-mermaid.md` | Complete import specification including redraw rules and edge-case handling |
| `skills/diagram-design/references/output-spec.md` | Dial configuration schemas and valid parameter values |
| `scripts/verify-mermaid-import.py` | Test harness for adversarial label validation and contract compliance |
| `scripts/fixtures/sample-adversarial.mmd` | Deliberately hostile test cases for security verification |

## Summary

- The **Mermaid import pipeline** enforces a strict trust boundary by treating all input as untrusted data and normalizing labels via `clean_label`.
- **Three stages** (Extract IR, Choose Dials, Redraw) transform raw Mermaid into editorial-quality diagrams without executing external content.
- **Multi-block Markdown** files process each fenced diagram independently, producing indexed output files.
- **Adversarial labels** containing HTML, scripts, or escaped characters are sanitized to plain text before reaching the rendering stage.
- The pipeline supports **budget analysis** (9 nodes, 12 edges) and generates **fidelity ledgers** documenting all editorial decisions.

## Frequently Asked Questions

### How does the pipeline handle multiple Mermaid diagrams in a single Markdown file?

The `mermaid_extract.py` script identifies each fenced Mermaid block (`` ```mermaid ``) through sequential scanning and processes them as independent diagram entities. When using `--diagram all`, the system generates separate output files with numeric suffixes (e.g., [`file-0.html`](https://github.com/cathrynlavery/diagram-design/blob/main/file-0.html), [`file-1.html`](https://github.com/cathrynlavery/diagram-design/blob/main/file-1.html)), ensuring no content is lost while maintaining isolation between diagrams.

### What security measures protect against malicious Mermaid code?

According to the cathrynlavery/diagram-design source code, the pipeline implements a strict trust boundary that never executes or renders raw Mermaid. The `clean_label` function strips HTML tags, entity-encodes special characters, removes `:::class` suffixes, and expands `@{ … }` attribute blocks before they enter the IR. All styling directives and clickable regions are discarded during pre-processing to prevent injection attacks.

### Can the pipeline convert Mermaid pie charts or mind maps?

No. The grammar detector in [`mermaid_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/mermaid_extract.py) explicitly supports only `flowchart`, `sequenceDiagram`, `stateDiagram-v2`, and `erDiagram` types. Unsupported diagram kinds trigger clear error messages directing users to supported alternatives, ensuring the editorial layout engine receives only compatible structural definitions.

### What is the "fidelity ledger" and why is it generated?

The fidelity ledger is a post-render report (e.g., "9 IR nodes → 6 drawn, 1 cycle collapsed") that documents every transformation applied during the redraw stage. This satisfies the anti-pattern policy that "no content is silently omitted" by explicitly logging merges, collapses, or drops, enabling auditors to verify that editorial simplifications were intentional rather than accidental data loss.