# How the Draw.io Import Extraction Process Works Internally in the Diagram-Design Repository

> Discover the draw.io import extraction process internally. Learn how this four-stage pipeline decodes, parses, analyzes, and renders draw.io files into an intermediate representation and Markdown report.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: internals
- Published: 2026-09-11

---

**The draw.io import extraction process follows a deterministic four-stage pipeline—payload decoding, XML parsing, structural analysis, and digest rendering—that converts any draw.io file format into a normalized intermediate representation (IR) and human-readable Markdown report.**

The `cathrynlavery/diagram-design` repository implements this pipeline in [`skills/diagram-design/scripts/drawio_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/drawio_extract.py), providing a stateless, format-agnostic importer that handles raw XML, PNG/SVG embeddings, and compressed payloads without making design decisions upstream.

## Stage 1: Payload Decoding and Format Detection

The extraction begins in `load_mxfile` (lines 62‑88), which acts as a central dispatcher to detect input formats and extract raw XML safely.

For **PNG files**, the `_png_embedded_xml` function (lines 13‑31) scans for the PNG magic bytes and extracts the embedded *mxfile* payload from metadata chunks. **SVG embeddings** are handled by `_svg_embedded_xml` (lines 52‑60), which parses the XML structure to locate the embedded diagram data.

When encountering **deflated base64 payloads**, the `_inflate` function (lines 93‑110) decompresses the stream using `_decompress_limited`, which enforces a hard `MAX_XML_BYTES` ceiling to prevent zip-bomb attacks. This safety-first approach ensures that malicious or malformed compressed inputs cannot exhaust system resources.

## Stage 2: XML Parsing and IR Construction

Once raw XML is extracted, `parse_file` (lines 23‑38) loads the document and determines whether it contains a single-page `<mxGraphModel>` or a multi-page `<mxfile>` wrapper.

The `parse_page` function (lines 58‑85) walks the DOM hierarchy in three distinct passes:

- **First pass**: Collects raw `<mxCell>` elements and converts them into `Node` dataclasses, capturing geometry, style strings, and labels using helper functions `parse_style`, `clean_label`, and `classify_shape` (lines 98‑140).
- **Second pass**: Resolves absolute positions and parent-child relationships, marking containers and computing nesting depth (lines 68‑75).
- **Third pass**: Creates `Edge` objects, processing waypoint arrays and attaching edge-labels (lines 76‑92).

The `shape_family` utility translates draw.io style strings into canonical shape identifiers, normalizing vendor-specific syntax into a consistent IR format.

## Stage 3: Structural Analysis and Heuristic Detection

After IR construction, the `analyze` function (lines 86‑136) computes structural signals used by downstream diagram-type selectors. This includes:

- Node and edge cardinality
- Maximum graph depth
- Shape frequency distributions
- Hub detection via centrality heuristics
- Candidate diagram classifications (sequence, flowchart, architecture)

Supporting utilities like `_has_cycle` (lines 45‑73) detect cyclic dependencies, while `_aligned` identifies lane-based layouts (such as swimlanes). These heuristics remain strictly descriptive—the extractor reports structural signals without conflating them with design recommendations.

## Stage 4: Digest Rendering and Output Serialization

The final stage transforms the IR into consumable outputs. The `digest` function (lines 113‑166) generates a human-readable Markdown report containing page-level statistics, node tables, and edge tables with proper Markdown escaping.

For programmatic consumption, `to_json` (lines 178‑235) emits the complete IR as JSON. The `main` function (lines 236‑262) wires these components to the CLI, parsing arguments and dispatching to the appropriate output formatter.

## Safety Mechanisms and Design Decisions

Three architectural decisions govern the extraction process:

- **Bounded Decompression**: The `_decompress_limited` implementation prevents decompression bombs by strictly enforcing `MAX_XML_BYTES` limits during inflate operations.
- **Format Agnosticism**: The `load_mxfile` dispatcher transparently handles raw XML, PNG, SVG, and URL-encoded payloads, eliminating preprocessing requirements for users.
- **Stateless Operation**: The script maintains strict separation between extraction and interpretation. The deterministic IR contains only structural facts, leaving diagram-type classification to higher-level logic such as the `import-drawio` command.

## Code Examples

You can invoke the extractor from the command line or import it programmatically.

**CLI usage—generate Markdown digest:**

```bash
python3 skills/diagram-design/scripts/drawio_extract.py my-diagram.drawio --page all

```

**CLI usage—export full IR as JSON:**

```bash
python3 skills/diagram-design/scripts/drawio_extract.py my-diagram.drawio --json > ir.json

```

**Programmatic usage in Python:**

```python
from pathlib import Path
from skills.diagram_design.scripts.drawio_extract import parse_file, digest, to_json

# Load draw.io file into IR objects

pages = parse_file(Path("my-diagram.drawio"))

# Generate human-readable report for first page

report = digest(Path("my-diagram.drawio"), pages, pages[:1], max_rows=30)
print(report)

# Export JSON representation

json_ir = to_json(Path("my-diagram.drawio"), pages, pages)
print(json_ir)

```

## Summary

- The draw.io import extraction process runs a four-stage pipeline: payload decoding, XML parsing, structural analysis, and digest rendering.
- `load_mxfile` handles multiple input formats (PNG, SVG, raw XML, compressed base64) with built-in zip-bomb protection via `_decompress_limited`.
- `parse_page` constructs a normalized IR using three-pass DOM walking to resolve geometry, hierarchy, and edge routing.
- The `analyze` function computes structural heuristics (cycles, hubs, layout alignment) without embedding design logic.
- Output formats include Markdown digests (`digest`) and complete JSON serialization (`to_json`), accessible via CLI or Python API.

## Frequently Asked Questions

### How does the extractor handle PNG or SVG files exported from draw.io?

The extractor detects PNG magic bytes via `_png_embedded_xml` and scans SVG structures via `_svg_embedded_xml` to locate embedded *mxfile* XML. Both functions extract the payload and pass it to the standard XML parsing pipeline, making the tool agnostic to how diagrams were exported.

### What prevents the extractor from crashing on malicious compressed inputs?

The `_decompress_limited` function enforces a hard `MAX_XML_BYTES` limit during inflation (implemented in `_inflate` at lines 93‑110). This bounds memory usage and prevents zip-bomb attacks that attempt to exploit decompression algorithms with exploding payload ratios.

### Can I use this extractor programmatically without invoking the CLI?

Yes. Import `parse_file`, `digest`, and `to_json` from `skills.diagram_design.scripts.drawio_extract` to integrate the pipeline into Python applications. The API accepts `pathlib.Path` objects and returns dataclass representations or formatted strings suitable for further processing.

### Why does the extractor produce an intermediate representation instead of final diagram types?

The architecture maintains strict separation between extraction and interpretation. By producing a stateless IR containing only structural facts (nodes, edges, geometry, styles), the system allows higher-level commands like `import-drawio` to apply evolving heuristics for diagram classification without modifying the parser.