How Diagram-Design Parses Excalidraw Import Scenes: A 10-Step Pipeline

Diagram-Design treats an Excalidraw scene as untrusted JSON and extracts a normalized intermediate representation (IR) through a rigorous parsing pipeline implemented in skills/diagram-design/scripts/excalidraw_extract.py.

The cathrynlavery/diagram-design repository converts Excalidraw drawings into structured diagram data through a multi-stage extraction process. Parsing Excalidraw import scenes requires validating file integrity, enforcing safety limits, and transforming arbitrary canvas elements into a deterministic node-edge model that downstream skills consume.

The 10-Stage Parsing Pipeline

The extractor processes every Excalidraw file through sequential validation and transformation stages.

1. File Validation

The extractor accepts only files ending in .excalidraw or .excalidraw.json. PNG and SVG exports are rejected immediately with a clear error message (lines 60-68).

2. Size and Content Limits

Input files larger than 16 MiB or scenes containing more than 10,000 elements are refused to prevent resource exhaustion (lines 75-94).

3. JSON Loading

The file is read as UTF-8 and parsed with json.loads. Any non-finite numbers (Infinity, NaN) cause immediate termination to ensure data integrity (lines 78-84).

4. Scene Sanity Checks

The top-level object must be a dictionary containing "type": "excalidraw" and an "elements" array. Missing or incorrect type fields abort the process (lines 87-92).

5. First Pass – Filtering and Binding

Deleted elements are counted and ignored. Text elements bound to other elements are collected in bound_labels for later folding into their parent shapes (lines 102-126).

6. Second Pass – Building Nodes

Elements are mapped to Node objects with type translation via NODE_SHAPES (e.g., "rectangle" → "rect"). Frames become container nodes, while image payloads, embeds, freedraw strokes, and unknown types are discarded but counted in statistics (lines 130-176). Frame membership is recorded with depth-1 parent-child links (lines 190-203).

7. Third Pass – Building Edges

Arrow and line elements become Edge objects. Arrowhead presence determines directionality (bidirectional, undirected). Bindings (startBinding, endBinding) resolve to node IDs, and waypoints are counted for path complexity (lines 209-240).

8. Degree Calculation

For each edge, source and target node degrees are updated. The system tracks in_degree, out_degree, and handles directed, undirected, and bidirectional edges appropriately (lines 244-266).

9. Structural Analysis

The analyze() function inspects the completed IR to compute statistics including node counts, edge counts, shape distributions, cycles, hubs, and entry points. It proposes candidate diagram types based on topological features (lines 370-482).

10. Output Generation

Depending on the --json flag, the script emits either the full IR as JSON via to_json() or a concise Markdown digest via digest() that downstream skills consume (lines 545-590). The import-excalidraw command orchestrates this extractor and forwards results to the diagram-generation pipeline (lines 40-49).

Working with the Excalidraw Extractor

Command-Line Usage

Generate a readable Markdown summary of an Excalidraw board:

python3 skills/diagram-design/scripts/excalidraw_extract.py \
    my-board.excalidraw

Export the full intermediate representation as JSON for programmatic processing:

python3 skills/diagram-design/scripts/excalidraw_extract.py \
    my-board.excalidraw --json > board.json

Programmatic Usage

Import the extractor functions directly from Python to integrate parsing into custom workflows:

from pathlib import Path
from skills.diagram_design.scripts.excalidraw_extract import (
    load_scene_data, parse_scene, to_json, digest
)

scene_path = Path("my-board.excalidraw")

# Validate and load raw JSON

doc = load_scene_data(scene_path)

# Build the intermediate representation

scene = parse_scene(scene_path, doc)

# Generate outputs

print(digest(scene_path, scene, max_rows=20))  # Markdown summary

print(to_json(scene_path, scene))              # JSON IR

Summary

  • File safety: Only .excalidraw and .excalidraw.json extensions are accepted, with strict 16 MiB and 10,000-element limits.
  • Validation pipeline: The parser rejects malformed JSON, non-finite numbers, and invalid scene types before processing.
  • Data sanitization: Images, embeds, freedraw strokes, and unknown types are discarded to ensure a clean node-edge model.
  • Structural intelligence: The analyze() function computes graph metrics and suggests diagram types based on topology.
  • Dual output: Extractor emits either detailed JSON IR or concise Markdown digests via command-line flags.

Frequently Asked Questions

What file formats does diagram-design accept for Excalidraw imports?

The extractor strictly accepts files ending in .excalidraw or .excalidraw.json. PNG and SVG exports are rejected with explicit error messages, as these formats require different parsing strategies not implemented in this pipeline.

How does diagram-design handle oversized Excalidraw scenes?

Files exceeding 16 MiB or containing more than 10,000 elements are refused immediately upon loading. These limits prevent memory exhaustion and ensure predictable processing times for the diagram-generation pipeline.

What happens to unsupported elements like images and freedraw strokes?

Image payloads, embeddable objects, freedraw strokes, and unrecognized element types are discarded during the second pass but counted in extraction statistics. This ensures the resulting graph contains only structural nodes and edges while maintaining awareness of omitted content.

How does the extractor determine edge directionality in Excalidraw scenes?

Directionality is determined by arrowhead presence on line and arrow elements. The parser checks for endArrowhead and startArrowhead properties to classify edges as bidirectional, undirected, or directed, updating in_degree and out_degree counts accordingly for topology analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →