How the draw.io Import Process Works in Diagram-Design
The draw.io import workflow in diagram-design follows a deterministic three-stage pipeline that extracts raw source files into a normalized intermediate representation, analyzes structural signals to select output parameters, and redraws the diagram from scratch using the design system's semantic model.
The draw.io import process in the cathrynlavery/diagram-design repository transforms raw .drawio, .drawio.png, .drawio.svg, or .xml files into clean, editorial-grade diagrams without preserving the original geometry, colors, or layout quirks. Unlike generic conversion tools, this pipeline treats source files as untrusted content and rebuilds diagrams semantically according to the design system's rules.
The Three-Stage Import Pipeline
The import mechanism operates through a strictly defined pipeline that ensures consistent, reproducible output regardless of the input file's compression method or complexity.
Stage 1: Extracting the Intermediate Representation (IR)
The process begins in skills/diagram-design/scripts/drawio_extract.py, which reads draw.io files without using generic Read tools because most inputs contain compressed payloads that appear as base-64 encoded data. The extractor detects the container type—whether raw XML, compressed <mxfile>, PNG-embedded mxfile, or SVG-embedded mxfile—and expands it safely while rejecting unsafe DTD or entity declarations.
The XML parses into a normalized intermediate representation (IR) consisting of Page → Node / Edge objects. This captures absolute positions, shape classifications, parent-child hierarchies, and structural signals such as hubs, cycles, and budget flags. The IR serves as the single source of truth for all subsequent processing stages.
Stage 2: Analysis and Dial Selection
Once extracted, the IR renders as a Markdown digest (or JSON via --json) that lists pages, node/edge counts, canvas size, shape distribution, type-candidates, budget warnings, and collapsible groups. The digest appears in the command output and drives the selection of four critical dials:
- format: html, svg, or png
- size: preset dimensions from
output-spec.md - detail: faithful, balanced, or simplified
- audience: engineer, mixed, or executive
The digest's budget: line triggers feasibility checks. For example, diagrams containing more than nine nodes cannot use the "faithful" detail level when targeting slide output. This logic resides in commands/import-drawio.md and skills/diagram-design/references/import-drawio.md.
Stage 3: Semantic Redrawing and Output Generation
Using the IR strictly as content reference, the skill constructs a fresh semantic model. It names the story, applies the chosen detail level by pruning via collapsible groups, selects focal hubs, rewrites labels for the target audience, and drops irrelevant edges.
Shapes map to the design system's treatments—for instance, cylinder converts to a Store/State box, while icon:aws becomes a monochrome cloud icon. Connectors reroute on a 4 px grid, discarding original waypoints. The system generates final HTML, optionally exports to SVG/PNG via the standard export pipeline, and attaches a fidelity ledger summarizing what was merged, collapsed, or dropped.
Key Implementation Files
The draw.io import process relies on specific files that define the extraction logic, command interface, and quality gates:
skills/diagram-design/scripts/drawio_extract.py: The core extractor that handles compression detection, safe XML parsing, and IR generation.skills/diagram-design/references/import-drawio.md: The procedural reference describing import steps, dial selection criteria, and redesign rules.commands/import-drawio.md: The slash-command definition that wires the extractor, analysis phase, and redraw logic together.scripts/verify-drawio-import.py: The CI verification harness ensuring the extractor, reference documentation, and command implementation remain synchronized.
Practical Usage Examples
Command-Line Extraction
Run the extractor directly to generate the IR and digest:
# Extract all pages from a draw.io file
python3 skills/diagram-design/scripts/drawio_extract.py \
diagrams/sample-architecture.drawio --page all
Full Import Command
Invoke the complete pipeline with explicit dials:
/diagram-design:import-drawio diagrams/sample-architecture.drawio \
--format html \
--size slide-16x9 \
--detail balanced \
--audience mixed
This command locates the skill directory, runs drawio_extract.py, parses the digest, selects the appropriate diagram type, applies the four dials, redraws the diagram from scratch, and writes output files while emitting a fidelity ledger.
Python API Integration
Extract and digest programmatically:
from pathlib import Path
from skills.diagram_design.scripts.drawio_extract import parse_file, digest
# Load a draw.io file and get the IR
pages = parse_file(Path("my-diagram.drawio"))
# Produce a human-readable digest (default 40 rows)
print(digest(Path("my-diagram.drawio"), pages, pages[:1], max_rows=40))
Access the JSON IR for downstream processing:
from skills.diagram_design.scripts.drawio_extract import parse_file, to_json
pages = parse_file(Path("my-diagram.drawio"))
json_ir = to_json(Path("my-diagram.drawio"), pages, pages[:1])
print(json_ir) # Full structure for downstream processing
Validation and Quality Assurance
Before writing any files, the skill runs the taste-gate defined in SKILL.md Section 9 and checks the checklist in output-spec.md Section 6. These gates ensure the output meets editorial standards for the selected audience and format.
The repository's CI pipeline executes scripts/verify-drawio-import.py to maintain integrity between the extractor logic, reference documentation, and command interface. This verification script ensures that changes to the extraction algorithm propagate correctly through the entire import chain.
Summary
- The draw.io import process treats source files as untrusted content, extracting only structural signals through a purpose-built parser in
drawio_extract.py. - The pipeline generates an intermediate representation (IR) that decouples input parsing from output generation, enabling flexible redraws.
- Four dials—format, size, detail, and audience—drive the semantic reconstruction phase, with feasibility checks based on node counts and canvas constraints.
- Connectors reroute on a 4 px grid and shapes map to the design system's treatments, ensuring brand consistency.
- Validation occurs through taste-gates and CI verification via
verify-drawio-import.py.
Frequently Asked Questions
What file formats does the draw.io import process support?
The import process supports .drawio, .drawio.png, .drawio.svg, and raw .xml files. The extractor in drawio_extract.py automatically detects whether the file contains raw XML, compressed <mxfile> payloads, or embedded mxfile data within PNG or SVG containers, then decompresses and parses them safely.
How does the system handle complex diagrams with many nodes?
When the IR analysis detects node counts exceeding specific thresholds—such as more than nine nodes for slide formats—the system enforces detail-level restrictions. The budget line in the Markdown digest triggers these feasibility checks, preventing users from selecting "faithful" detail for outputs where physical constraints would compromise readability.
Can I use the draw.io importer programmatically without the slash command?
Yes. Import parse_file, digest, and to_json directly from skills.diagram_design.scripts.drawio_extract to work with the intermediate representation in Python. This API allows custom workflows that bypass the standard slash command while still leveraging the extraction and normalization logic.
What is the fidelity ledger and where can I find it?
The fidelity ledger is a metadata attachment generated during Stage 3 that documents all transformations applied to the source diagram. It lists which elements were merged, collapsed, or dropped during the redraw phase, providing transparency about the semantic distance between the original draw.io file and the final output. The ledger accompanies the generated HTML or exported image files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →