# Understanding the draw.io Import Pipeline and Compressed Payload Handling

> Explore the draw.io import pipeline and compressed payload handling. Learn how draw.io securely processes files with decompression limits and zip bomb prevention.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: deep-dive
- Published: 2026-09-09

---

**The draw.io import pipeline converts raw .drawio, .drawio.png, and .drawio.svg files into a structured intermediate representation through four stages—container unpacking, payload decompression, IR building, and digest export—while enforcing 64 MiB decompression limits and blocking unsafe XML entities to prevent zip bomb attacks.**

The `cathrynlavery/diagram-design` repository implements a robust draw.io import pipeline that transforms proprietary diagram formats into clean, design-system-compliant outputs. This pipeline handles multiple container types including plain XML, deflate-compressed Base64 payloads, and embedded PNG or SVG metadata. Understanding how this draw.io import pipeline processes compressed payloads is essential for developers integrating diagram conversion into automated workflows.

## The Four Stages of the draw.io Import Pipeline

The pipeline processes every input file through four logical stages, each implemented in [`skills/diagram-design/scripts/drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/drawio_extract.py).

### Stage 1: Load and Unpack Containers

The `load_mxfile()` function detects the container format and returns raw `<mxfile>` or `<mxGraphModel>` XML. The pipeline supports four container variants:

- Plain XML files
- Deflate-compressed Base64 strings (standard .drawio format)
- PNG files with embedded `mxfile` chunks
- SVG files with diagram data in the `content` attribute

The function dispatches to `_png_embedded_xml()`, `_svg_embedded_xml()`, or `_inflate()` based on file signatures and content headers.

### Stage 2: Decode Compressed Payloads

Many draw.io files store diagrams as **deflate-compressed Base64** that is also URL-encoded. The `_inflate()` function orchestrates decompression, calling `_decompress_limited()` to enforce a strict **64 MiB ceiling** on output size. This protects against zip bombs by streaming decompression and checking the accumulator after each chunk. If the limit is exceeded, the pipeline raises `PayloadTooLarge` and exits with code 2.

### Stage 3: Build the Intermediate Representation

Once decoded, the XML flows through `parse_file()` → `parse_page()` → `analyze()`. This chain resolves absolute geometry, records nodes and edges, identifies containers, and computes structural signals including hub detection, cycle identification, and collapsible group candidates. The resulting IR captures the semantic structure of the diagram independent of its original coordinates or styling.

### Stage 4: Export Digest and JSON

The final stage emits either a human-readable Markdown digest via `digest()` or a full JSON intermediate representation via `to_json()`. These outputs feed downstream commands like `/diagram-design:import-drawio`, which apply the "four dials" (format, size, detail, audience) from [`output-spec.md`](https://github.com/cathrynlavery/diagram-design/blob/main/output-spec.md) to regenerate the diagram in the Diagram Design visual system.

## Deep Dive into Compressed Payload Handling

The draw.io import pipeline implements multiple defensive layers when handling compressed payloads, as detailed in [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py).

### Base64 Decoding and Multiple wbits Support

The decoder first applies `base64.b64decode(payload)` at line 96. Because draw.io uses raw deflate streams without standard zlib headers, the code tests three window bits values (`-15`, `15`, and `47`) at lines 99-103 to support both raw deflate and zlib-wrapped streams. This brute-force approach ensures compatibility with various export settings from the draw.io editor.

### Size-Bounded Decompression Security

The `_decompress_limited()` function (lines 70-81) implements streaming decompression with an explicit size counter. After each decompression chunk, the function checks whether adding the next segment would exceed the 64 MiB limit. If so, it aborts immediately with exit code 2 and prints `drawio_extract: PayloadTooLarge`. This guarantees memory safety when processing untrusted user uploads.

### URL Decoding and XML Safety Checks

After inflation, the text may still be URL-encoded; `unquote(text)` at line 108 removes percent-escapes. Before parsing, `_reject_unsafe_xml()` (lines 62-66) scans for `<!DOCTYPE>` or `<!ENTITY>` declarations that could trigger external entity expansion attacks. Any detected unsafe XML causes immediate termination with exit code 2, ensuring the pipeline never processes malicious DTD references.

## Integration with the Diagram Design Skill

The skill's command reference in [`commands/import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/commands/import-drawio.md) triggers the pipeline when detecting `.drawio*` files. The workflow proceeds as follows:

1. **Extraction**: The skill launches `python3 <skill-dir>/scripts/drawio_extract.py` against the input file
2. **Type Selection**: The digest provides node/edge tables and type candidates, allowing the skill to select appropriate `type-*.md` references (Architecture, Flowchart, etc.)
3. **Semantic Modeling**: The IR enables story extraction, focal node identification, and label rewriting
4. **Redraw**: A fresh layout generates on a 4 px grid, discarding source coordinates, colors, and shapes per Step 5 of [`import-drawio.md`](https://github.com/cathrynlavery/diagram-design/blob/main/import-drawio.md)
5. **Verification**: Changes to the import path must pass [`scripts/verify-drawio-import.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-drawio-import.py), which validates the extractor against `sample-architecture.drawio` and ensures reference consistency

## Code Examples

### Import a draw.io file via the skill command

```bash
/diagram-design:import-drawio platform.drawio \
  --size=slide-16x9 --detail=simplified --audience=executive

```

*The command parses `platform.drawio` with [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py), selects the Architecture type, and produces an HTML file matching a 16 : 9 slide.*

### Run the extractor directly for debugging

```bash
python3 skills/diagram-design/scripts/drawio_extract.py \
  scripts/fixtures/sample-architecture.drawio \
  --page=all \
  --max-rows=20 \
  --out /tmp/arch-digest.md

```

*Produces a Markdown digest for every page in the sample file, limited to 20 rows per table.*

### Extract the full JSON IR programmatically

```bash
python3 skills/diagram-design/scripts/drawio_extract.py \
  diagrams/example.drawio \
  --json \
  --out /tmp/example.json

```

*The JSON output contains the complete node/edge list for custom tooling or analysis.*

### Handle a PNG-embedded draw.io diagram

```bash
python3 skills/diagram-design/scripts/drawio_extract.py \
  images/diagram.png \
  --out /tmp/png-digest.md

```

*The extractor automatically locates and extracts the embedded `mxfile` chunk from the PNG metadata.*

## Summary

- The draw.io import pipeline consists of four stages—Load, Decode, Build IR, and Export—implemented primarily in [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py)
- Compressed payload handling uses `base64.b64decode()`, tests multiple `wbits` values (-15, 15, 47), and enforces a 64 MiB ceiling via `_decompress_limited()`
- Security measures include zip-bomb protection and `_reject_unsafe_xml()` filtering to block DTD/ENTITY declarations
- The pipeline exits with code 2 on any extraction failure, preventing silent fallbacks to raw file reading
- Output formats include Markdown digests for humans and JSON IR for programmatic consumption

## Frequently Asked Questions

### How does the draw.io import pipeline handle different container formats?

The `load_mxfile()` function in [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py) detects container types by examining file headers and extensions. For PNG files, it calls `_png_embedded_xml()` to extract the `mxfile` chunk from PNG metadata. For SVG files, `_svg_embedded_xml()` retrieves diagram data from the `content` attribute. Plain XML files pass through directly, while compressed payloads route to `_inflate()`.

### What security measures protect against zip bombs in compressed payloads?

The pipeline implements `_decompress_limited()` which streams decompression in chunks while maintaining a running byte counter. If decompressed output would exceed 64 MiB, the function raises `PayloadTooLarge` and the process exits with code 2. Additionally, `_reject_unsafe_xml()` blocks documents containing `<!DOCTYPE>` or `<!ENTITY>` declarations before parsing begins.

### How can I extract the intermediate representation programmatically?

Use the `--json` flag when invoking [`drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/drawio_extract.py) directly. This calls `to_json()` instead of `digest()`, emitting the complete intermediate representation including node geometries, edge connections, and structural signals. The JSON format is suitable for integration with external layout engines or automated diagram analysis tools.

### What happens when the import pipeline encounters unsafe XML?

Before parsing any extracted XML, the pipeline invokes `_reject_unsafe_xml()` (lines 62-66) to scan for DTD or ENTITY declarations. If detected, the extractor prints a descriptive error message prefixed with `drawio_extract:` and terminates with exit code 2. This prevents XXE (XML External Entity) attacks and ensures malicious payloads never reach the parsing stage.