# How to Extract a Dominant Palette and Font Stack from a Website for Diagram Branding

> Automate brand extraction for diagram design. Learn how to get a dominant palette and font stack from any website to ensure consistent, accessible branding with cathrynlavery/diagram-design.

- Repository: [Cathryn Lavery/diagram-design](https://github.com/cathrynlavery/diagram-design)
- Tags: how-to-guide
- Published: 2026-09-08

---

**TLDR:** The `cathrynlavery/diagram-design` repository automates brand extraction by fetching a homepage, parsing HTML and CSS for color and typography values, mapping them to semantic tokens like `paper` and `ink`, validating WCAG AA contrast, and persisting the results to [`style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/style-guide.md) via the `font_families()` and `extract_colors()` helpers in [`scripts/verify-docs-sync.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-docs-sync.py).

Extracting a dominant palette and font stack from a website for diagram branding is the core capability of the diagram-design onboarding flow. By analyzing live CSS and DOM structures, the tool converts any URL into a cohesive set of design tokens in under a minute. This ensures generated diagrams inherit the exact visual identity of the target site, from background colors to heading typefaces.

## How the Onboarding Pipeline Extracts Brand Assets

### Fetching and Parsing the Target Homepage

The agent initiates a standard HTTP GET request to capture the raw HTML. It then scans the DOM for specific CSS properties: the `<body>` background becomes the `paper` token, primary text color maps to `ink`, secondary captions become `muted`, major container backgrounds resolve to `paper-2`, and the most frequently used brand color (typically CTAs or links) is assigned to `accent`.

### Typography Extraction from CSS and Google Fonts

Font discovery targets three key selectors. Heading fonts from `<h1>` elements map to the `title` token, body text fonts populate `node-name`, and monospaced faces from `<code>` or `<pre>` blocks become `sublabel`. The `font_families(url: str) -> set[str]` function in [`scripts/verify-docs-sync.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-docs-sync.py) specifically hunts for `<link href="…fonts.googleapis.com…">` tags while walking the DOM to build the complete stack.

### Semantic Token Mapping

Raw hex values and font names are translated into design-system semantics defined in [`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md). This mapping table, reproduced in [`README.md`](https://github.com/cathrynlavery/diagram-design/blob/main/README.md) (lines 13–22), serves as the single source of truth for all diagram templates, ensuring consistency across exports.

### Automated Contrast Validation

Before persisting, the system validates WCAG AA compliance for `ink` versus `paper` at the 9–12 px font sizes used for diagram labels. If the native site colors fail accessibility thresholds, the pipeline proposes an adjusted value, as documented in [`README.md`](https://github.com/cathrynlavery/diagram-design/blob/main/README.md) (lines 26–28).

## Programmatic Extraction with Python Helpers

For debugging or custom integrations, you can invoke the extraction logic directly from [`scripts/verify-docs-sync.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-docs-sync.py). The module exposes two primary functions: `font_families()` for typography and `extract_colors()` for the palette.

```python

# Run inside the repo checkout

from scripts.verify_docs_sync import font_families, extract_colors

url = "https://example.com"
fonts = font_families(url)               # → {'Geist', 'Instrument Serif', ...}

colors = extract_colors(url)             # → {'paper': '#f5f5f5', 'ink': '#2d3142', ...}

print("Fonts:", fonts)
print("Colors:", colors)

```

These helpers perform the same DOM walking and CSS variable resolution (e.g., `--color-paper`) used by the natural-language interface.

## Natural-Language vs. Manual Workflows

The repository supports two distinct modes for extracting brand assets.

**Natural-language onboarding** (recommended for standard use) requires a single prompt:

```text
You:  "onboard diagram-design to https://example.com"

Agent: "Fetched https://example.com – extracted palette:
  paper:#f5f5f5  ink:#2d3142  muted:#4f5d75  accent:#eb6c36
  fonts: 'Geist', 'Geist Mono', 'Instrument Serif'
Apply?"

```

**Manual extraction** via Python (demonstrated above) is ideal for CI/CD pipelines, batch processing multiple URLs, or verifying extraction logic during development.

## Persistent Storage and Confirmation

Once extracted, the agent presents a diff of generated tokens and awaits confirmation. Upon approval, values are written to [`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md). The [`README.md`](https://github.com/cathrynlavery/diagram-design/blob/main/README.md) (line 99) describes this final persistence step, which completes the branding setup and enables all subsequent diagrams to automatically inherit the captured `Instrument Serif` titles, `Geist` node names, and `Geist Mono` sub-labels.

## Summary

- The onboarding flow in `cathrynlavery/diagram-design` fetches a URL and parses HTML/CSS to extract a dominant palette and font stack from a website for diagram branding.
- Extraction logic resides in [`scripts/verify-docs-sync.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-docs-sync.py), specifically the `font_families()` and `extract_colors()` functions.
- Colors map to semantic tokens: `paper`, `ink`, `muted`, `paper-2`, and `accent`.
- Fonts categorize into `title` (headings), `node-name` (body), and `sublabel` (monospace).
- WCAG AA contrast validation ensures readability at 9–12 px label sizes.
- Final tokens persist to [`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md) after user confirmation.

## Frequently Asked Questions

### What semantic tokens does diagram-design extract from a website?

The pipeline extracts five color tokens—`paper` (background), `ink` (primary text), `muted` (captions), `paper-2` (containers), and `accent` (brand/CTA)—plus three font tokens: `title` from `<h1>` elements, `node-name` from body text, and `sublabel` from `<code>` or `<pre>` blocks.

### How does the contrast validation work for diagram labels?

Before saving tokens, the system checks WCAG AA compliance specifically for the `ink` color against the `paper` background at 9–12 px font sizes. If the combination fails accessibility standards, the pipeline suggests an adjusted color value to ensure diagram readability.

### Can I manually extract palette and font data without using the natural-language interface?

Yes. Import `font_families` and `extract_colors` from [`scripts/verify-docs-sync.py`](https://github.com/cathrynlavery/diagram-design/blob/main/scripts/verify-docs-sync.py) to programmatically extract data from any URL. This approach is documented in [`skills/diagram-design/references/onboarding.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/onboarding.md) and supports batch processing or debugging.

### Where are the extracted brand tokens stored in the repository?

Confirmed tokens are written to [`skills/diagram-design/references/style-guide.md`](https://github.com/cathrynlavery/diagram-design/blob/main/skills/diagram-design/references/style-guide.md), which serves as the persistent single source of truth for all diagram templates. The mapping specifications are also summarized in [`README.md`](https://github.com/cathrynlavery/diagram-design/blob/main/README.md) lines 13–22.