Data Flow Diagram Capabilities in Archify: Schema, Validation, and Rendering Guide
Archify's Data Flow renderer enables deterministic visualization of data pipelines, lineage, PII handling, and consumer relationships through a strict JSON schema that compiles to annotated SVG diagrams.
The tt-a1i/archify repository provides a schema-driven diagramming toolkit where Data Flow diagrams serve as the dedicated renderer for mapping how data moves through complex systems. Unlike generic drawing tools, Archify enforces structural constraints through archify/schemas/dataflow.schema.json, ensuring that every diagram accurately represents sources, transforms, stores, and consumption points while highlighting sensitivity boundaries.
Core Purpose and Use Cases
According to the source code in README.md, Data Flow diagrams describe the movement of data through a system, making sources, transforms, stores, and consumption points explicit. The primary capabilities include:
- Pipeline visualization – Map analytics pipelines such as product-analytics flows
- Lineage tracking – Trace data from origin through transformation to consumption
- PII handling – Highlight privacy boundaries and sensitive data paths
- Consumer relationships – Identify downstream dependencies and data stores
These diagrams communicate data-sensitivity and governance rules to both technical and non-technical stakeholders, as implemented in the repository's gallery examples.
Schema Structure and Required Fields
All Data Flow diagrams must conform to archify/schemas/dataflow.schema.json, which enforces a typed intermediate representation (IR) with specific structural requirements.
Fixed Diagram Type and Top-Level Keys
The schema requires diagram_type: "dataflow" as an immutable constant. Every valid document must include four required top-level keys:
meta– Metadata including title, visual preset, and quality profilestages– Ordered processing phases (maximum five)nodes– System components (services, storage, external systems)flows– Directed connections between nodes
Stage Definitions
The stages array supports up to five ordered processing stages that organize nodes horizontally. A typical configuration follows the pattern:
- Sources – Data origination points
- Ingest – Collection and intake services
- Process – Transformation and stream processing
- Store – Persistent storage systems
- Consume – Analytics and visualization endpoints
Each stage object requires a label field as defined in the schema lines 100-105.
Node Specifications
Nodes model architectural components with five required fields:
id– Unique identifier referenced by flowstype– Component classification (client,service,security,database,consumer)label– Human-readable display namestage– Integer index (0-4) mapping to the stages arrayrow– Vertical positioning within a stage
The schema defines these requirements in lines 19-33 of archify/schemas/dataflow.schema.json.
Flow Connections and Classifications
Flows define directed edges between nodes using three required fields:
from– Source node IDto– Target node IDlabel– Descriptive text explaining the data movement
Optional visual modifiers control rendering:
classification– Applies semantic styling (emphasis,security)emphasis– Boolean flag for primary hot-path highlightingsecurity– Boolean flag for PII/privacy boundariesdashed– Indicates batch or asynchronous derived paths
These style flags drive color and line-style rendering in the final SVG output.
Visual Features and Legend System
The renderer automatically generates a legend that distinguishes visual elements based on the legend object defined in the schema (lines 66-80). The legend identifies:
- Emphasis – Primary hot-path data flows
- Security – PII or privacy-boundary crossings
- Dashed – Batch or asynchronous processing paths
- Database – Storage node types
These classifications ensure that sensitivity boundaries and critical paths remain visually distinct in the generated artifact.
Quality Profiles for Publication Readiness
Archify implements two validation profiles when processing Data Flow diagrams:
standard– Basic schema compliance and structural validationshowcase– Enforces stricter layout rules including single SVG output, orthogonal arrows, and legend clearance requirements
The showcase profile ensures publication-ready artifacts suitable for documentation and stakeholder presentations, as referenced in README.md lines 145-146.
Implementation Workflow
Creating a Data Flow diagram follows a three-step deterministic process:
- Author – Create a typed JSON IR file describing your pipeline structure
- Validate – Run
archify validate dataflow --json <file>to check schema compliance and visual constraints - Render – Execute
archify deliver dataflow <input> <output.html>to generate a deterministic HTML/SVG artifact
Complete Example
The following JSON defines a product analytics pipeline with PII handling:
{
"schema_version": 1,
"diagram_type": "dataflow",
"meta": {
"title": "Product Analytics Data Flow",
"visual_preset": "signal-flow",
"quality_profile": "showcase"
},
"stages": [
{ "label": "Sources" },
{ "label": "Ingest" },
{ "label": "Process" },
{ "label": "Store" },
{ "label": "Consume" }
],
"nodes": [
{ "id": "web", "type": "client", "label": "Web App", "stage": 0, "row": 0 },
{ "id": "edge", "type": "service", "label": "Edge API", "stage": 1, "row": 0 },
{ "id": "stream", "type": "service", "label": "Event Stream", "stage": 2, "row": 0 },
{ "id": "vault", "type": "security", "label": "PII Vault", "stage": 3, "row": 0 },
{ "id": "warehouse", "type": "database", "label": "Data Warehouse", "stage": 3, "row": 1 },
{ "id": "dashboard", "type": "consumer", "label": "Dashboard", "stage": 4, "row": 0 }
],
"flows": [
{ "from": "web", "to": "edge", "label": "click → API", "classification": "emphasis" },
{ "from": "edge", "to": "stream", "label": "event publish", "classification": "emphasis" },
{ "from": "stream", "to": "vault", "label": "PII → vault", "classification": "security" },
{ "from": "stream", "to": "warehouse", "label": "analytics → warehouse", "classification": "emphasis" },
{ "from": "warehouse", "to": "dashboard", "label": "report data", "classification": "emphasis" }
]
}
Save this as my-dataflow.json, then execute:
npx archify.mjs validate dataflow my-dataflow.json --json
npx archify.mjs deliver dataflow my-dataflow.json ./my-dataflow.html --open
The validator checks against archify/schemas/dataflow.schema.json, applies the showcase quality profile, and outputs a deterministic HTML file containing the annotated SVG diagram.
Summary
- Data Flow diagrams in Archify provide schema-driven visualization of data pipelines, lineage, and PII boundaries through deterministic JSON-to-SVG rendering.
- The
archify/schemas/dataflow.schema.jsonfile enforces required structures includingmeta,stages,nodes, andflows, with a maximum of five ordered processing stages. - Visual annotations automatically apply through the legend system to highlight emphasis paths, security boundaries, database storage, and asynchronous flows.
- Quality profiles (
standardandshowcase) control validation strictness, withshowcaseensuring publication-ready artifacts with orthogonal arrows and proper legend clearance. - The
archify.mjsCLI providesvalidateanddelivercommands for checking schema compliance and generating shareable HTML/SVG outputs.
Frequently Asked Questions
What file defines the structure for Data Flow diagrams in Archify?
The formal JSON Schema resides in archify/schemas/dataflow.schema.json. This file defines the required diagram_type: "dataflow" constant, mandatory top-level keys (meta, stages, nodes, flows), and validation rules for node fields and flow classifications.
How many processing stages can a Data Flow diagram include?
Archify supports up to five ordered processing stages in the stages array. Common patterns include Sources → Ingest → Process → Store → Consume, though you may define fewer stages as needed for your specific pipeline architecture.
What CLI commands validate and render Data Flow diagrams?
Use archify validate dataflow <file> --json to check schema compliance and visual constraints against archify/schemas/dataflow.schema.json. Then use archify deliver dataflow <input> <output.html> to generate the deterministic SVG artifact. Both commands are implemented in archify/bin/archify.mjs.
How does Archify indicate PII or security boundaries in diagrams?
Set the classification: "security" flag on flow objects in your JSON definition. The renderer applies the security visual style (defined in the schema's legend object) to highlight PII movements and privacy boundaries, making sensitive data paths immediately visible to reviewers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →