How to Extract a Dominant Palette and Font Stack from a Website for Diagram Branding

TLDR: The cathrynlavery/diagram-design repository automates brand extraction by fetching a homepage, parsing HTML and CSS for color and typography values, mapping them to semantic tokens like paper and ink, validating WCAG AA contrast, and persisting the results to style-guide.md via the font_families() and extract_colors() helpers in scripts/verify-docs-sync.py.

Extracting a dominant palette and font stack from a website for diagram branding is the core capability of the diagram-design onboarding flow. By analyzing live CSS and DOM structures, the tool converts any URL into a cohesive set of design tokens in under a minute. This ensures generated diagrams inherit the exact visual identity of the target site, from background colors to heading typefaces.

How the Onboarding Pipeline Extracts Brand Assets

Fetching and Parsing the Target Homepage

The agent initiates a standard HTTP GET request to capture the raw HTML. It then scans the DOM for specific CSS properties: the <body> background becomes the paper token, primary text color maps to ink, secondary captions become muted, major container backgrounds resolve to paper-2, and the most frequently used brand color (typically CTAs or links) is assigned to accent.

Typography Extraction from CSS and Google Fonts

Font discovery targets three key selectors. Heading fonts from <h1> elements map to the title token, body text fonts populate node-name, and monospaced faces from <code> or <pre> blocks become sublabel. The font_families(url: str) -> set[str] function in scripts/verify-docs-sync.py specifically hunts for <link href="…fonts.googleapis.com…"> tags while walking the DOM to build the complete stack.

Semantic Token Mapping

Raw hex values and font names are translated into design-system semantics defined in skills/diagram-design/references/style-guide.md. This mapping table, reproduced in README.md (lines 13–22), serves as the single source of truth for all diagram templates, ensuring consistency across exports.

Automated Contrast Validation

Before persisting, the system validates WCAG AA compliance for ink versus paper at the 9–12 px font sizes used for diagram labels. If the native site colors fail accessibility thresholds, the pipeline proposes an adjusted value, as documented in README.md (lines 26–28).

Programmatic Extraction with Python Helpers

For debugging or custom integrations, you can invoke the extraction logic directly from scripts/verify-docs-sync.py. The module exposes two primary functions: font_families() for typography and extract_colors() for the palette.


# Run inside the repo checkout

from scripts.verify_docs_sync import font_families, extract_colors

url = "https://example.com"
fonts = font_families(url)               # → {'Geist', 'Instrument Serif', ...}

colors = extract_colors(url)             # → {'paper': '#f5f5f5', 'ink': '#2d3142', ...}

print("Fonts:", fonts)
print("Colors:", colors)

These helpers perform the same DOM walking and CSS variable resolution (e.g., --color-paper) used by the natural-language interface.

Natural-Language vs. Manual Workflows

The repository supports two distinct modes for extracting brand assets.

Natural-language onboarding (recommended for standard use) requires a single prompt:

You:  "onboard diagram-design to https://example.com"

Agent: "Fetched https://example.com – extracted palette:
  paper:#f5f5f5  ink:#2d3142  muted:#4f5d75  accent:#eb6c36
  fonts: 'Geist', 'Geist Mono', 'Instrument Serif'
Apply?"

Manual extraction via Python (demonstrated above) is ideal for CI/CD pipelines, batch processing multiple URLs, or verifying extraction logic during development.

Persistent Storage and Confirmation

Once extracted, the agent presents a diff of generated tokens and awaits confirmation. Upon approval, values are written to skills/diagram-design/references/style-guide.md. The README.md (line 99) describes this final persistence step, which completes the branding setup and enables all subsequent diagrams to automatically inherit the captured Instrument Serif titles, Geist node names, and Geist Mono sub-labels.

Summary

  • The onboarding flow in cathrynlavery/diagram-design fetches a URL and parses HTML/CSS to extract a dominant palette and font stack from a website for diagram branding.
  • Extraction logic resides in scripts/verify-docs-sync.py, specifically the font_families() and extract_colors() functions.
  • Colors map to semantic tokens: paper, ink, muted, paper-2, and accent.
  • Fonts categorize into title (headings), node-name (body), and sublabel (monospace).
  • WCAG AA contrast validation ensures readability at 9–12 px label sizes.
  • Final tokens persist to skills/diagram-design/references/style-guide.md after user confirmation.

Frequently Asked Questions

What semantic tokens does diagram-design extract from a website?

The pipeline extracts five color tokens—paper (background), ink (primary text), muted (captions), paper-2 (containers), and accent (brand/CTA)—plus three font tokens: title from <h1> elements, node-name from body text, and sublabel from <code> or <pre> blocks.

How does the contrast validation work for diagram labels?

Before saving tokens, the system checks WCAG AA compliance specifically for the ink color against the paper background at 9–12 px font sizes. If the combination fails accessibility standards, the pipeline suggests an adjusted color value to ensure diagram readability.

Can I manually extract palette and font data without using the natural-language interface?

Yes. Import font_families and extract_colors from scripts/verify-docs-sync.py to programmatically extract data from any URL. This approach is documented in skills/diagram-design/references/onboarding.md and supports batch processing or debugging.

Where are the extracted brand tokens stored in the repository?

Confirmed tokens are written to skills/diagram-design/references/style-guide.md, which serves as the persistent single source of truth for all diagram templates. The mapping specifications are also summarized in README.md lines 13–22.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →