# How to Onboard Brand Color Extraction in Diagram-Design: A Complete Technical Guide

> Learn how Diagram Design automates brand color extraction by scraping CSS and parsing design tokens. This guide shows you how to onboard brand colors for seamless adoption.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Diagram‑Design automates brand adoption by scraping CSS, parsing JSON design tokens, and heuristically mapping external colors to internal semantic roles like `accent` and `ink`, then persists the palette to `style‑guide.md`.**

The `cathrynlavery/diagram-design` repository provides a systematic technique for extracting and standardizing brand colors from diverse sources. This onboarding process eliminates manual color picking by analyzing websites, local design systems, or installed skills to generate a consistent visual identity for every diagram.

## The Onboarding Pipeline

The extraction technique follows a deterministic pipeline defined in [[`skills/diagram-design/references/onboarding.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/onboarding.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/onboarding.md). The system accepts multiple input types and normalizes them into a unified token structure.

### Source Detection and Input Methods

The engine first prompts the user to specify how to obtain the brand palette. Valid sources include:

- A live **website URL** for CSS scraping
- An installed **`brand-design`** or **`ui-kit`** skill
- A local **design-system folder** containing token files
- A **manual token list** provided by the user

### File Discovery and Prioritization

When processing a local folder, the engine scans for candidate files in the root directory. It prioritizes filenames containing `color`, `token`, `brand`, `palette`, `style`, or `theme`. If the scan returns more than 20 files, the system narrows the selection to the most relevant candidates based on naming conventions and file extensions.

### Color Extraction Strategies

The system employs two distinct parsing strategies depending on the source format.

**HTML/CSS Scraping**

For website URLs, the scraper analyzes the live CSS to identify the most-used brand colors. It specifically targets visual elements that typically carry brand identity:

- **CTA buttons** (primary action colors)
- **Link elements** (interactive text colors)
- **Heading accents** (decorative typography colors)

The dominant color discovered in these selectors is mapped to the `accent` token.

**JSON Token Parsing**

For structured data, the parser recognizes **Style Dictionary** format objects (e.g., `{ "color": { "brand": { "value": "#eb6c36" } } }`). The engine flattens the nested paths and applies heuristics to leaf keys. Keys matching `brand`, `primary`, `cta`, or `highlight` are automatically mapped to the internal `accent` role.

### Heuristic Mapping to Semantic Tokens

Raw extracted values undergo heuristic processing to assign semantic meaning. The engine maintains a mapping dictionary that translates external naming conventions to Diagram-Design's internal taxonomy:

- `primary`, `cta`, `highlight` → `accent`
- `background`, `surface` → `paper`
- `text`, `body` → `ink`

This mapping ensures that brand colors align with the semantic roles defined in the style system.

### Palette Limiting and Receipt Generation

To maintain visual consistency, the engine caps the final palette to **three core colors**: `paper`, `ink`, and `accent`. When the extraction yields more than 20 raw colors, only the top three most frequent or brand-like values are retained.

After processing, the system generates a **brand fidelity receipt**. This report documents which external tokens were mapped to internal roles and is required whenever the user explicitly requests brand-matched output. The receipt attaches to the preview to verify extraction accuracy.

## Persisting the Brand Profile

Once extraction completes, the final color scheme is written to [[`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md). This file serves as the single source of truth, defining semantic roles that the rendering engine consumes to skin every diagram automatically.

Users can persist the extracted palette as a named profile (e.g., `.diagram-design` marker with `profile: <slug>`). Subsequent diagrams in that project reuse the saved profile without re-extraction, referencing the stored style guide directly.

## Implementation Examples

The following examples demonstrate how to invoke the onboarding pipeline programmatically and via CLI.

```python

# Extract brand palette from a website URL

from pathlib import Path
from diagram_design.onboard import extract_brand_palette

url = "https://example.com"
palette = extract_brand_palette(url)

# Returns dict with semantic mapping

print(palette)

# {'paper': '#f5f5f5', 'ink': '#2d3142', 'accent': '#eb6c36'}

# Load a previously saved profile

from diagram_design.profile import load_profile

profile = load_profile("my-brand")
print(profile.style_guide_path)

# PosixPath('~/.diagram-design/profiles/my-brand/style-guide.md')

```

```bash

# CLI onboarding workflow

diagram-design onboard \
  --source url https://mycompany.com \
  --save-profile my-brand

# Output: Extracts colors, writes style-guide.md, saves profile

```

## Core Source Files and Architecture

Understanding the onboarding technique requires familiarity with these key files:

| File | Role |
|------|------|
| [[`skills/diagram-design/references/onboarding.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/onboarding.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/onboarding.md) | Defines the complete extraction flow, source handling, heuristics, and receipt generation logic. |
| [[`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md) | Central token definition file that receives the extracted brand palette and defines semantic roles. |
| [[`skills/diagram-design/SKILL.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/SKILL.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/SKILL.md) | Documents the skill's capabilities, including the onboarding prompt and rules against silently shipping default-skinned diagrams. |
| [[`commands/profile.md`](https://github.com/cathrynlavery/diagram-design/blob/main/commands/profile.md)](https://github.com/cathrynlavery/diagram-design/blob/main/commands/profile.md) | CLI documentation for saving, loading, and switching between brand profiles. |

## Summary

- **Diagram-Design** extracts brand colors via CSS scraping or JSON parsing, mapping external values to semantic tokens like `accent` and `ink`.
- The onboarding pipeline prioritizes CTA buttons, links, and headings when scraping websites, and recognizes Style Dictionary formats for JSON inputs.
- Heuristic mapping translates common token names (primary, brand, cta) to internal roles, limiting the final palette to three colors for consistency.
- Extraction results are persisted to [`style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/style-guide.md) and can be saved as named profiles for reuse across projects.

## Frequently Asked Questions

### How does Diagram-Design extract colors from a website without an API?

The system scrapes the target URL's CSS to identify the most frequently used brand colors appearing in CTA buttons, link elements, and heading accents. According to the source code in [`onboarding.md`](https://github.com/cathrynlavery/diagram-design/blob/main/onboarding.md), it prioritizes these values as the `accent` token during the extraction process.

### What happens if my design system has more than three brand colors?

The engine limits the palette to the three most critical semantic roles—`paper`, `ink`, and `accent`—to maintain visual consistency. When more than 20 raw colors are detected, only the top three most frequent or brand-like values are retained based on the heuristics implemented in the repository.

### Can I use existing Style Dictionary JSON files for onboarding?

Yes. The parser recognizes the Style Dictionary format (`{ "color": { "brand": { "value": "#..." } } }`), flattens the token paths, and maps keys such as `primary`, `cta`, or `highlight` to internal roles like `accent` and `brand` as defined in the onboarding specifications.

### Where is the extracted palette stored after onboarding?

The final color scheme is written to [[`style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/style-guide.md)](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md), which serves as the single source of truth for all diagram rendering in the project. Users can also save the configuration as a named profile for future use.