Understanding the draw.io Import Pipeline and Compressed Payload Handling
The draw.io import pipeline converts raw .drawio, .drawio.png, and .drawio.svg files into a structured intermediate representation through four stages—container unpacking, payload decompression, IR building, and digest export—while enforcing 64 MiB decompression limits and blocking unsafe XML entities to prevent zip bomb attacks.
The cathrynlavery/diagram-design repository implements a robust draw.io import pipeline that transforms proprietary diagram formats into clean, design-system-compliant outputs. This pipeline handles multiple container types including plain XML, deflate-compressed Base64 payloads, and embedded PNG or SVG metadata. Understanding how this draw.io import pipeline processes compressed payloads is essential for developers integrating diagram conversion into automated workflows.
The Four Stages of the draw.io Import Pipeline
The pipeline processes every input file through four logical stages, each implemented in skills/diagram-design/scripts/drawio_extract.py.
Stage 1: Load and Unpack Containers
The load_mxfile() function detects the container format and returns raw <mxfile> or <mxGraphModel> XML. The pipeline supports four container variants:
- Plain XML files
- Deflate-compressed Base64 strings (standard .drawio format)
- PNG files with embedded
mxfilechunks - SVG files with diagram data in the
contentattribute
The function dispatches to _png_embedded_xml(), _svg_embedded_xml(), or _inflate() based on file signatures and content headers.
Stage 2: Decode Compressed Payloads
Many draw.io files store diagrams as deflate-compressed Base64 that is also URL-encoded. The _inflate() function orchestrates decompression, calling _decompress_limited() to enforce a strict 64 MiB ceiling on output size. This protects against zip bombs by streaming decompression and checking the accumulator after each chunk. If the limit is exceeded, the pipeline raises PayloadTooLarge and exits with code 2.
Stage 3: Build the Intermediate Representation
Once decoded, the XML flows through parse_file() → parse_page() → analyze(). This chain resolves absolute geometry, records nodes and edges, identifies containers, and computes structural signals including hub detection, cycle identification, and collapsible group candidates. The resulting IR captures the semantic structure of the diagram independent of its original coordinates or styling.
Stage 4: Export Digest and JSON
The final stage emits either a human-readable Markdown digest via digest() or a full JSON intermediate representation via to_json(). These outputs feed downstream commands like /diagram-design:import-drawio, which apply the "four dials" (format, size, detail, audience) from output-spec.md to regenerate the diagram in the Diagram Design visual system.
Deep Dive into Compressed Payload Handling
The draw.io import pipeline implements multiple defensive layers when handling compressed payloads, as detailed in drawio_extract.py.
Base64 Decoding and Multiple wbits Support
The decoder first applies base64.b64decode(payload) at line 96. Because draw.io uses raw deflate streams without standard zlib headers, the code tests three window bits values (-15, 15, and 47) at lines 99-103 to support both raw deflate and zlib-wrapped streams. This brute-force approach ensures compatibility with various export settings from the draw.io editor.
Size-Bounded Decompression Security
The _decompress_limited() function (lines 70-81) implements streaming decompression with an explicit size counter. After each decompression chunk, the function checks whether adding the next segment would exceed the 64 MiB limit. If so, it aborts immediately with exit code 2 and prints drawio_extract: PayloadTooLarge. This guarantees memory safety when processing untrusted user uploads.
URL Decoding and XML Safety Checks
After inflation, the text may still be URL-encoded; unquote(text) at line 108 removes percent-escapes. Before parsing, _reject_unsafe_xml() (lines 62-66) scans for <!DOCTYPE> or <!ENTITY> declarations that could trigger external entity expansion attacks. Any detected unsafe XML causes immediate termination with exit code 2, ensuring the pipeline never processes malicious DTD references.
Integration with the Diagram Design Skill
The skill's command reference in commands/import-drawio.md triggers the pipeline when detecting .drawio* files. The workflow proceeds as follows:
- Extraction: The skill launches
python3 <skill-dir>/scripts/drawio_extract.pyagainst the input file - Type Selection: The digest provides node/edge tables and type candidates, allowing the skill to select appropriate
type-*.mdreferences (Architecture, Flowchart, etc.) - Semantic Modeling: The IR enables story extraction, focal node identification, and label rewriting
- Redraw: A fresh layout generates on a 4 px grid, discarding source coordinates, colors, and shapes per Step 5 of
import-drawio.md - Verification: Changes to the import path must pass
scripts/verify-drawio-import.py, which validates the extractor againstsample-architecture.drawioand ensures reference consistency
Code Examples
Import a draw.io file via the skill command
/diagram-design:import-drawio platform.drawio \
--size=slide-16x9 --detail=simplified --audience=executive
The command parses platform.drawio with drawio_extract.py, selects the Architecture type, and produces an HTML file matching a 16 : 9 slide.
Run the extractor directly for debugging
python3 skills/diagram-design/scripts/drawio_extract.py \
scripts/fixtures/sample-architecture.drawio \
--page=all \
--max-rows=20 \
--out /tmp/arch-digest.md
Produces a Markdown digest for every page in the sample file, limited to 20 rows per table.
Extract the full JSON IR programmatically
python3 skills/diagram-design/scripts/drawio_extract.py \
diagrams/example.drawio \
--json \
--out /tmp/example.json
The JSON output contains the complete node/edge list for custom tooling or analysis.
Handle a PNG-embedded draw.io diagram
python3 skills/diagram-design/scripts/drawio_extract.py \
images/diagram.png \
--out /tmp/png-digest.md
The extractor automatically locates and extracts the embedded mxfile chunk from the PNG metadata.
Summary
- The draw.io import pipeline consists of four stages—Load, Decode, Build IR, and Export—implemented primarily in
drawio_extract.py - Compressed payload handling uses
base64.b64decode(), tests multiplewbitsvalues (-15, 15, 47), and enforces a 64 MiB ceiling via_decompress_limited() - Security measures include zip-bomb protection and
_reject_unsafe_xml()filtering to block DTD/ENTITY declarations - The pipeline exits with code 2 on any extraction failure, preventing silent fallbacks to raw file reading
- Output formats include Markdown digests for humans and JSON IR for programmatic consumption
Frequently Asked Questions
How does the draw.io import pipeline handle different container formats?
The load_mxfile() function in drawio_extract.py detects container types by examining file headers and extensions. For PNG files, it calls _png_embedded_xml() to extract the mxfile chunk from PNG metadata. For SVG files, _svg_embedded_xml() retrieves diagram data from the content attribute. Plain XML files pass through directly, while compressed payloads route to _inflate().
What security measures protect against zip bombs in compressed payloads?
The pipeline implements _decompress_limited() which streams decompression in chunks while maintaining a running byte counter. If decompressed output would exceed 64 MiB, the function raises PayloadTooLarge and the process exits with code 2. Additionally, _reject_unsafe_xml() blocks documents containing <!DOCTYPE> or <!ENTITY> declarations before parsing begins.
How can I extract the intermediate representation programmatically?
Use the --json flag when invoking drawio_extract.py directly. This calls to_json() instead of digest(), emitting the complete intermediate representation including node geometries, edge connections, and structural signals. The JSON format is suitable for integration with external layout engines or automated diagram analysis tools.
What happens when the import pipeline encounters unsafe XML?
Before parsing any extracted XML, the pipeline invokes _reject_unsafe_xml() (lines 62-66) to scan for DTD or ENTITY declarations. If detected, the extractor prints a descriptive error message prefixed with drawio_extract: and terminates with exit code 2. This prevents XXE (XML External Entity) attacks and ensures malicious payloads never reach the parsing stage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →