How the Draw.io Import Extraction Process Works Internally in the Diagram-Design Repository
The draw.io import extraction process follows a deterministic four-stage pipeline—payload decoding, XML parsing, structural analysis, and digest rendering—that converts any draw.io file format into a normalized intermediate representation (IR) and human-readable Markdown report.
The cathrynlavery/diagram-design repository implements this pipeline in skills/diagram-design/scripts/drawio_extract.py, providing a stateless, format-agnostic importer that handles raw XML, PNG/SVG embeddings, and compressed payloads without making design decisions upstream.
Stage 1: Payload Decoding and Format Detection
The extraction begins in load_mxfile (lines 62‑88), which acts as a central dispatcher to detect input formats and extract raw XML safely.
For PNG files, the _png_embedded_xml function (lines 13‑31) scans for the PNG magic bytes and extracts the embedded mxfile payload from metadata chunks. SVG embeddings are handled by _svg_embedded_xml (lines 52‑60), which parses the XML structure to locate the embedded diagram data.
When encountering deflated base64 payloads, the _inflate function (lines 93‑110) decompresses the stream using _decompress_limited, which enforces a hard MAX_XML_BYTES ceiling to prevent zip-bomb attacks. This safety-first approach ensures that malicious or malformed compressed inputs cannot exhaust system resources.
Stage 2: XML Parsing and IR Construction
Once raw XML is extracted, parse_file (lines 23‑38) loads the document and determines whether it contains a single-page <mxGraphModel> or a multi-page <mxfile> wrapper.
The parse_page function (lines 58‑85) walks the DOM hierarchy in three distinct passes:
- First pass: Collects raw
<mxCell>elements and converts them intoNodedataclasses, capturing geometry, style strings, and labels using helper functionsparse_style,clean_label, andclassify_shape(lines 98‑140). - Second pass: Resolves absolute positions and parent-child relationships, marking containers and computing nesting depth (lines 68‑75).
- Third pass: Creates
Edgeobjects, processing waypoint arrays and attaching edge-labels (lines 76‑92).
The shape_family utility translates draw.io style strings into canonical shape identifiers, normalizing vendor-specific syntax into a consistent IR format.
Stage 3: Structural Analysis and Heuristic Detection
After IR construction, the analyze function (lines 86‑136) computes structural signals used by downstream diagram-type selectors. This includes:
- Node and edge cardinality
- Maximum graph depth
- Shape frequency distributions
- Hub detection via centrality heuristics
- Candidate diagram classifications (sequence, flowchart, architecture)
Supporting utilities like _has_cycle (lines 45‑73) detect cyclic dependencies, while _aligned identifies lane-based layouts (such as swimlanes). These heuristics remain strictly descriptive—the extractor reports structural signals without conflating them with design recommendations.
Stage 4: Digest Rendering and Output Serialization
The final stage transforms the IR into consumable outputs. The digest function (lines 113‑166) generates a human-readable Markdown report containing page-level statistics, node tables, and edge tables with proper Markdown escaping.
For programmatic consumption, to_json (lines 178‑235) emits the complete IR as JSON. The main function (lines 236‑262) wires these components to the CLI, parsing arguments and dispatching to the appropriate output formatter.
Safety Mechanisms and Design Decisions
Three architectural decisions govern the extraction process:
- Bounded Decompression: The
_decompress_limitedimplementation prevents decompression bombs by strictly enforcingMAX_XML_BYTESlimits during inflate operations. - Format Agnosticism: The
load_mxfiledispatcher transparently handles raw XML, PNG, SVG, and URL-encoded payloads, eliminating preprocessing requirements for users. - Stateless Operation: The script maintains strict separation between extraction and interpretation. The deterministic IR contains only structural facts, leaving diagram-type classification to higher-level logic such as the
import-drawiocommand.
Code Examples
You can invoke the extractor from the command line or import it programmatically.
CLI usage—generate Markdown digest:
python3 skills/diagram-design/scripts/drawio_extract.py my-diagram.drawio --page all
CLI usage—export full IR as JSON:
python3 skills/diagram-design/scripts/drawio_extract.py my-diagram.drawio --json > ir.json
Programmatic usage in Python:
from pathlib import Path
from skills.diagram_design.scripts.drawio_extract import parse_file, digest, to_json
# Load draw.io file into IR objects
pages = parse_file(Path("my-diagram.drawio"))
# Generate human-readable report for first page
report = digest(Path("my-diagram.drawio"), pages, pages[:1], max_rows=30)
print(report)
# Export JSON representation
json_ir = to_json(Path("my-diagram.drawio"), pages, pages)
print(json_ir)
Summary
- The draw.io import extraction process runs a four-stage pipeline: payload decoding, XML parsing, structural analysis, and digest rendering.
load_mxfilehandles multiple input formats (PNG, SVG, raw XML, compressed base64) with built-in zip-bomb protection via_decompress_limited.parse_pageconstructs a normalized IR using three-pass DOM walking to resolve geometry, hierarchy, and edge routing.- The
analyzefunction computes structural heuristics (cycles, hubs, layout alignment) without embedding design logic. - Output formats include Markdown digests (
digest) and complete JSON serialization (to_json), accessible via CLI or Python API.
Frequently Asked Questions
How does the extractor handle PNG or SVG files exported from draw.io?
The extractor detects PNG magic bytes via _png_embedded_xml and scans SVG structures via _svg_embedded_xml to locate embedded mxfile XML. Both functions extract the payload and pass it to the standard XML parsing pipeline, making the tool agnostic to how diagrams were exported.
What prevents the extractor from crashing on malicious compressed inputs?
The _decompress_limited function enforces a hard MAX_XML_BYTES limit during inflation (implemented in _inflate at lines 93‑110). This bounds memory usage and prevents zip-bomb attacks that attempt to exploit decompression algorithms with exploding payload ratios.
Can I use this extractor programmatically without invoking the CLI?
Yes. Import parse_file, digest, and to_json from skills.diagram_design.scripts.drawio_extract to integrate the pipeline into Python applications. The API accepts pathlib.Path objects and returns dataclass representations or formatted strings suitable for further processing.
Why does the extractor produce an intermediate representation instead of final diagram types?
The architecture maintains strict separation between extraction and interpretation. By producing a stateless IR containing only structural facts (nodes, edges, geometry, styles), the system allows higher-level commands like import-drawio to apply evolving heuristics for diagram classification without modifying the parser.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →