How the draw.io Import Process Works in Diagram-Design

The draw.io import workflow in diagram-design follows a deterministic three-stage pipeline that extracts raw source files into a normalized intermediate representation, analyzes structural signals to select output parameters, and redraws the diagram from scratch using the design system's semantic model.

The draw.io import process in the cathrynlavery/diagram-design repository transforms raw .drawio, .drawio.png, .drawio.svg, or .xml files into clean, editorial-grade diagrams without preserving the original geometry, colors, or layout quirks. Unlike generic conversion tools, this pipeline treats source files as untrusted content and rebuilds diagrams semantically according to the design system's rules.

The Three-Stage Import Pipeline

The import mechanism operates through a strictly defined pipeline that ensures consistent, reproducible output regardless of the input file's compression method or complexity.

Stage 1: Extracting the Intermediate Representation (IR)

The process begins in skills/diagram-design/scripts/drawio_extract.py, which reads draw.io files without using generic Read tools because most inputs contain compressed payloads that appear as base-64 encoded data. The extractor detects the container type—whether raw XML, compressed <mxfile>, PNG-embedded mxfile, or SVG-embedded mxfile—and expands it safely while rejecting unsafe DTD or entity declarations.

The XML parses into a normalized intermediate representation (IR) consisting of Page → Node / Edge objects. This captures absolute positions, shape classifications, parent-child hierarchies, and structural signals such as hubs, cycles, and budget flags. The IR serves as the single source of truth for all subsequent processing stages.

Stage 2: Analysis and Dial Selection

Once extracted, the IR renders as a Markdown digest (or JSON via --json) that lists pages, node/edge counts, canvas size, shape distribution, type-candidates, budget warnings, and collapsible groups. The digest appears in the command output and drives the selection of four critical dials:

  • format: html, svg, or png
  • size: preset dimensions from output-spec.md
  • detail: faithful, balanced, or simplified
  • audience: engineer, mixed, or executive

The digest's budget: line triggers feasibility checks. For example, diagrams containing more than nine nodes cannot use the "faithful" detail level when targeting slide output. This logic resides in commands/import-drawio.md and skills/diagram-design/references/import-drawio.md.

Stage 3: Semantic Redrawing and Output Generation

Using the IR strictly as content reference, the skill constructs a fresh semantic model. It names the story, applies the chosen detail level by pruning via collapsible groups, selects focal hubs, rewrites labels for the target audience, and drops irrelevant edges.

Shapes map to the design system's treatments—for instance, cylinder converts to a Store/State box, while icon:aws becomes a monochrome cloud icon. Connectors reroute on a 4 px grid, discarding original waypoints. The system generates final HTML, optionally exports to SVG/PNG via the standard export pipeline, and attaches a fidelity ledger summarizing what was merged, collapsed, or dropped.

Key Implementation Files

The draw.io import process relies on specific files that define the extraction logic, command interface, and quality gates:

Practical Usage Examples

Command-Line Extraction

Run the extractor directly to generate the IR and digest:


# Extract all pages from a draw.io file

python3 skills/diagram-design/scripts/drawio_extract.py \
    diagrams/sample-architecture.drawio --page all

Full Import Command

Invoke the complete pipeline with explicit dials:

/diagram-design:import-drawio diagrams/sample-architecture.drawio \
    --format html \
    --size slide-16x9 \
    --detail balanced \
    --audience mixed

This command locates the skill directory, runs drawio_extract.py, parses the digest, selects the appropriate diagram type, applies the four dials, redraws the diagram from scratch, and writes output files while emitting a fidelity ledger.

Python API Integration

Extract and digest programmatically:

from pathlib import Path
from skills.diagram_design.scripts.drawio_extract import parse_file, digest

# Load a draw.io file and get the IR

pages = parse_file(Path("my-diagram.drawio"))

# Produce a human-readable digest (default 40 rows)

print(digest(Path("my-diagram.drawio"), pages, pages[:1], max_rows=40))

Access the JSON IR for downstream processing:

from skills.diagram_design.scripts.drawio_extract import parse_file, to_json

pages = parse_file(Path("my-diagram.drawio"))
json_ir = to_json(Path("my-diagram.drawio"), pages, pages[:1])
print(json_ir)  # Full structure for downstream processing

Validation and Quality Assurance

Before writing any files, the skill runs the taste-gate defined in SKILL.md Section 9 and checks the checklist in output-spec.md Section 6. These gates ensure the output meets editorial standards for the selected audience and format.

The repository's CI pipeline executes scripts/verify-drawio-import.py to maintain integrity between the extractor logic, reference documentation, and command interface. This verification script ensures that changes to the extraction algorithm propagate correctly through the entire import chain.

Summary

  • The draw.io import process treats source files as untrusted content, extracting only structural signals through a purpose-built parser in drawio_extract.py.
  • The pipeline generates an intermediate representation (IR) that decouples input parsing from output generation, enabling flexible redraws.
  • Four dials—format, size, detail, and audience—drive the semantic reconstruction phase, with feasibility checks based on node counts and canvas constraints.
  • Connectors reroute on a 4 px grid and shapes map to the design system's treatments, ensuring brand consistency.
  • Validation occurs through taste-gates and CI verification via verify-drawio-import.py.

Frequently Asked Questions

What file formats does the draw.io import process support?

The import process supports .drawio, .drawio.png, .drawio.svg, and raw .xml files. The extractor in drawio_extract.py automatically detects whether the file contains raw XML, compressed <mxfile> payloads, or embedded mxfile data within PNG or SVG containers, then decompresses and parses them safely.

How does the system handle complex diagrams with many nodes?

When the IR analysis detects node counts exceeding specific thresholds—such as more than nine nodes for slide formats—the system enforces detail-level restrictions. The budget line in the Markdown digest triggers these feasibility checks, preventing users from selecting "faithful" detail for outputs where physical constraints would compromise readability.

Can I use the draw.io importer programmatically without the slash command?

Yes. Import parse_file, digest, and to_json directly from skills.diagram_design.scripts.drawio_extract to work with the intermediate representation in Python. This API allows custom workflows that bypass the standard slash command while still leveraging the extraction and normalization logic.

What is the fidelity ledger and where can I find it?

The fidelity ledger is a metadata attachment generated during Stage 3 that documents all transformations applied to the source diagram. It lists which elements were merged, collapsed, or dropped during the redraw phase, providing transparency about the semantic distance between the original draw.io file and the final output. The ledger accompanies the generated HTML or exported image files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →