# How the draw.io Import Process Works in Diagram-Design

> Understand the drawio import process in diagram-design. Discover the three-stage pipeline that normalizes, analyzes, and redraws diagrams for a seamless workflow.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: internals
- Published: 2026-09-13

---

**The draw.io import workflow in diagram-design follows a deterministic three-stage pipeline that extracts raw source files into a normalized intermediate representation, analyzes structural signals to select output parameters, and redraws the diagram from scratch using the design system's semantic model.**

The **draw.io import process** in the `cathrynlavery/diagram-design` repository transforms raw `.drawio`, `.drawio.png`, `.drawio.svg`, or `.xml` files into clean, editorial-grade diagrams without preserving the original geometry, colors, or layout quirks. Unlike generic conversion tools, this pipeline treats source files as untrusted content and rebuilds diagrams semantically according to the design system's rules.

## The Three-Stage Import Pipeline

The import mechanism operates through a strictly defined pipeline that ensures consistent, reproducible output regardless of the input file's compression method or complexity.

### Stage 1: Extracting the Intermediate Representation (IR)

The process begins in [`skills/diagram-design/scripts/drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/drawio_extract.py), which reads draw.io files without using generic `Read` tools because most inputs contain compressed payloads that appear as base-64 encoded data. The extractor detects the container type—whether raw XML, compressed `<mxfile>`, PNG-embedded mxfile, or SVG-embedded mxfile—and expands it safely while rejecting unsafe DTD or entity declarations.

The XML parses into a **normalized intermediate representation (IR)** consisting of `Page → Node / Edge` objects. This captures absolute positions, shape classifications, parent-child hierarchies, and structural signals such as hubs, cycles, and budget flags. The IR serves as the single source of truth for all subsequent processing stages.

### Stage 2: Analysis and Dial Selection

Once extracted, the IR renders as a **Markdown digest** (or JSON via `--json`) that lists pages, node/edge counts, canvas size, shape distribution, type-candidates, budget warnings, and collapsible groups. The digest appears in the command output and drives the selection of four critical **dials**:

- **format**: html, svg, or png
- **size**: preset dimensions from [`output-spec.md`](https://github.com/cathrynlavery/diagram-design/blob/main/output-spec.md)
- **detail**: faithful, balanced, or simplified
- **audience**: engineer, mixed, or executive

The digest's `budget:` line triggers feasibility checks. For example, diagrams containing more than nine nodes cannot use the "faithful" detail level when targeting slide output. This logic resides in [`commands/import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/commands/import-drawio.md) and [`skills/diagram-design/references/import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/import-drawio.md).

### Stage 3: Semantic Redrawing and Output Generation

Using the IR strictly as content reference, the skill constructs a fresh semantic model. It names the story, applies the chosen detail level by pruning via collapsible groups, selects focal hubs, rewrites labels for the target audience, and drops irrelevant edges. 

Shapes map to the design system's treatments—for instance, `cylinder` converts to a Store/State box, while `icon:aws` becomes a monochrome cloud icon. Connectors reroute on a 4 px grid, discarding original waypoints. The system generates final HTML, optionally exports to SVG/PNG via the standard export pipeline, and attaches a **fidelity ledger** summarizing what was merged, collapsed, or dropped.

## Key Implementation Files

The draw.io import process relies on specific files that define the extraction logic, command interface, and quality gates:

- **[`skills/diagram-design/scripts/drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/drawio_extract.py)**: The core extractor that handles compression detection, safe XML parsing, and IR generation.
- **[`skills/diagram-design/references/import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/import-drawio.md)**: The procedural reference describing import steps, dial selection criteria, and redesign rules.
- **[`commands/import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/commands/import-drawio.md)**: The slash-command definition that wires the extractor, analysis phase, and redraw logic together.
- **[`scripts/verify-drawio-import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-drawio-import.py)**: The CI verification harness ensuring the extractor, reference documentation, and command implementation remain synchronized.

## Practical Usage Examples

### Command-Line Extraction

Run the extractor directly to generate the IR and digest:

```bash

# Extract all pages from a draw.io file

python3 skills/diagram-design/scripts/drawio_extract.py \
    diagrams/sample-architecture.drawio --page all

```

### Full Import Command

Invoke the complete pipeline with explicit dials:

```bash
/diagram-design:import-drawio diagrams/sample-architecture.drawio \
    --format html \
    --size slide-16x9 \
    --detail balanced \
    --audience mixed

```

This command locates the skill directory, runs [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py), parses the digest, selects the appropriate diagram type, applies the four dials, redraws the diagram from scratch, and writes output files while emitting a fidelity ledger.

### Python API Integration

Extract and digest programmatically:

```python
from pathlib import Path
from skills.diagram_design.scripts.drawio_extract import parse_file, digest

# Load a draw.io file and get the IR

pages = parse_file(Path("my-diagram.drawio"))

# Produce a human-readable digest (default 40 rows)

print(digest(Path("my-diagram.drawio"), pages, pages[:1], max_rows=40))

```

Access the JSON IR for downstream processing:

```python
from skills.diagram_design.scripts.drawio_extract import parse_file, to_json

pages = parse_file(Path("my-diagram.drawio"))
json_ir = to_json(Path("my-diagram.drawio"), pages, pages[:1])
print(json_ir)  # Full structure for downstream processing

```

## Validation and Quality Assurance

Before writing any files, the skill runs the **taste-gate** defined in [`SKILL.md`](https://github.com/cathrynlavery/diagram-design/blob/main/SKILL.md) Section 9 and checks the checklist in [`output-spec.md`](https://github.com/cathrynlavery/diagram-design/blob/main/output-spec.md) Section 6. These gates ensure the output meets editorial standards for the selected audience and format.

The repository's CI pipeline executes [`scripts/verify-drawio-import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-drawio-import.py) to maintain integrity between the extractor logic, reference documentation, and command interface. This verification script ensures that changes to the extraction algorithm propagate correctly through the entire import chain.

## Summary

- The **draw.io import process** treats source files as untrusted content, extracting only structural signals through a purpose-built parser in [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py).
- The pipeline generates an **intermediate representation (IR)** that decouples input parsing from output generation, enabling flexible redraws.
- Four **dials**—format, size, detail, and audience—drive the semantic reconstruction phase, with feasibility checks based on node counts and canvas constraints.
- Connectors reroute on a 4 px grid and shapes map to the design system's treatments, ensuring brand consistency.
- **Validation** occurs through taste-gates and CI verification via [`verify-drawio-import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/verify-drawio-import.py).

## Frequently Asked Questions

### What file formats does the draw.io import process support?

The import process supports `.drawio`, `.drawio.png`, `.drawio.svg`, and raw `.xml` files. The extractor in [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py) automatically detects whether the file contains raw XML, compressed `<mxfile>` payloads, or embedded mxfile data within PNG or SVG containers, then decompresses and parses them safely.

### How does the system handle complex diagrams with many nodes?

When the IR analysis detects node counts exceeding specific thresholds—such as more than nine nodes for slide formats—the system enforces detail-level restrictions. The **budget** line in the Markdown digest triggers these feasibility checks, preventing users from selecting "faithful" detail for outputs where physical constraints would compromise readability.

### Can I use the draw.io importer programmatically without the slash command?

Yes. Import `parse_file`, `digest`, and `to_json` directly from `skills.diagram_design.scripts.drawio_extract` to work with the intermediate representation in Python. This API allows custom workflows that bypass the standard slash command while still leveraging the extraction and normalization logic.

### What is the fidelity ledger and where can I find it?

The **fidelity ledger** is a metadata attachment generated during Stage 3 that documents all transformations applied to the source diagram. It lists which elements were merged, collapsed, or dropped during the redraw phase, providing transparency about the semantic distance between the original draw.io file and the final output. The ledger accompanies the generated HTML or exported image files.