# How Diagram-Design Parses Excalidraw Import Scenes: A 10-Step Pipeline

> Discover how Diagram-Design parses Excalidraw import scenes through a 10-step pipeline. Learn about the untrusted JSON treatment and intermediate representation extraction process.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: internals
- Published: 2026-09-11

---

**Diagram-Design treats an Excalidraw scene as untrusted JSON and extracts a normalized intermediate representation (IR) through a rigorous parsing pipeline implemented in [`skills/diagram-design/scripts/excalidraw_extract.py`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py).**

The `cathrynlavery/diagram-design` repository converts Excalidraw drawings into structured diagram data through a multi-stage extraction process. Parsing Excalidraw import scenes requires validating file integrity, enforcing safety limits, and transforming arbitrary canvas elements into a deterministic node-edge model that downstream skills consume.

## The 10-Stage Parsing Pipeline

The extractor processes every Excalidraw file through sequential validation and transformation stages.

### 1. File Validation

The extractor accepts only files ending in `.excalidraw` or [`.excalidraw.json`](https://github.com/cathrynlavery/diagram-design/blob/main/.excalidraw.json). PNG and SVG exports are rejected immediately with a clear error message ([lines 60-68](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L60-L68)).

### 2. Size and Content Limits

Input files larger than **16 MiB** or scenes containing more than **10,000 elements** are refused to prevent resource exhaustion ([lines 75-94](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L75-L94)).

### 3. JSON Loading

The file is read as UTF-8 and parsed with `json.loads`. Any non-finite numbers (`Infinity`, `NaN`) cause immediate termination to ensure data integrity ([lines 78-84](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L78-L84)).

### 4. Scene Sanity Checks

The top-level object must be a dictionary containing `"type": "excalidraw"` and an `"elements"` array. Missing or incorrect type fields abort the process ([lines 87-92](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L87-L92)).

### 5. First Pass – Filtering and Binding

Deleted elements are counted and ignored. Text elements bound to other elements are collected in `bound_labels` for later folding into their parent shapes ([lines 102-126](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L102-L126)).

### 6. Second Pass – Building Nodes

Elements are mapped to **`Node`** objects with type translation via `NODE_SHAPES` (e.g., `"rectangle"` → `"rect"`). Frames become container nodes, while image payloads, embeds, freedraw strokes, and unknown types are discarded but counted in statistics ([lines 130-176](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L130-L176)). Frame membership is recorded with depth-1 parent-child links ([lines 190-203](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L190-L203)).

### 7. Third Pass – Building Edges

Arrow and line elements become **`Edge`** objects. Arrowhead presence determines directionality (`bidirectional`, `undirected`). Bindings (`startBinding`, `endBinding`) resolve to node IDs, and waypoints are counted for path complexity ([lines 209-240](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L209-L240)).

### 8. Degree Calculation

For each edge, source and target node degrees are updated. The system tracks `in_degree`, `out_degree`, and handles directed, undirected, and bidirectional edges appropriately ([lines 244-266](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L244-L266)).

### 9. Structural Analysis

The `analyze()` function inspects the completed IR to compute statistics including node counts, edge counts, shape distributions, cycles, hubs, and entry points. It proposes candidate diagram types based on topological features ([lines 370-482](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L370-L482)).

### 10. Output Generation

Depending on the `--json` flag, the script emits either the full IR as JSON via `to_json()` or a concise Markdown digest via `digest()` that downstream skills consume ([lines 545-590](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/scripts/excalidraw_extract.py#L545-L590)). The `import-excalidraw` command orchestrates this extractor and forwards results to the diagram-generation pipeline ([lines 40-49](https://github.com/cathrynlavery/diagram-design/blob/main/commands/import-excalidraw.md#L40-L49)).

## Working with the Excalidraw Extractor

### Command-Line Usage

Generate a readable Markdown summary of an Excalidraw board:

```bash
python3 skills/diagram-design/scripts/excalidraw_extract.py \
    my-board.excalidraw

```

Export the full intermediate representation as JSON for programmatic processing:

```bash
python3 skills/diagram-design/scripts/excalidraw_extract.py \
    my-board.excalidraw --json > board.json

```

### Programmatic Usage

Import the extractor functions directly from Python to integrate parsing into custom workflows:

```python
from pathlib import Path
from skills.diagram_design.scripts.excalidraw_extract import (
    load_scene_data, parse_scene, to_json, digest
)

scene_path = Path("my-board.excalidraw")

# Validate and load raw JSON

doc = load_scene_data(scene_path)

# Build the intermediate representation

scene = parse_scene(scene_path, doc)

# Generate outputs

print(digest(scene_path, scene, max_rows=20))  # Markdown summary

print(to_json(scene_path, scene))              # JSON IR

```

## Summary

- **File safety**: Only `.excalidraw` and [`.excalidraw.json`](https://github.com/cathrynlavery/diagram-design/blob/main/.excalidraw.json) extensions are accepted, with strict 16 MiB and 10,000-element limits.
- **Validation pipeline**: The parser rejects malformed JSON, non-finite numbers, and invalid scene types before processing.
- **Data sanitization**: Images, embeds, freedraw strokes, and unknown types are discarded to ensure a clean node-edge model.
- **Structural intelligence**: The `analyze()` function computes graph metrics and suggests diagram types based on topology.
- **Dual output**: Extractor emits either detailed JSON IR or concise Markdown digests via command-line flags.

## Frequently Asked Questions

### What file formats does diagram-design accept for Excalidraw imports?

The extractor strictly accepts files ending in `.excalidraw` or [`.excalidraw.json`](https://github.com/cathrynlavery/diagram-design/blob/main/.excalidraw.json). PNG and SVG exports are rejected with explicit error messages, as these formats require different parsing strategies not implemented in this pipeline.

### How does diagram-design handle oversized Excalidraw scenes?

Files exceeding **16 MiB** or containing more than **10,000 elements** are refused immediately upon loading. These limits prevent memory exhaustion and ensure predictable processing times for the diagram-generation pipeline.

### What happens to unsupported elements like images and freedraw strokes?

Image payloads, embeddable objects, freedraw strokes, and unrecognized element types are discarded during the second pass but counted in extraction statistics. This ensures the resulting graph contains only structural nodes and edges while maintaining awareness of omitted content.

### How does the extractor determine edge directionality in Excalidraw scenes?

Directionality is determined by arrowhead presence on line and arrow elements. The parser checks for `endArrowhead` and `startArrowhead` properties to classify edges as `bidirectional`, `undirected`, or directed, updating `in_degree` and `out_degree` counts accordingly for topology analysis.