How the AI Website Cloner Template Extracts Design Tokens from a Target Site
The ai-website-cloner-template extracts design tokens by delegating to AI agents that use headless browser automation to inspect computed CSS values, aggregate them into semantic token files, and generate Tailwind v4 configurations.
The ai-website-cloner-template repository provides an AI-driven workflow for cloning websites by capturing their visual design systems. Rather than embedding static extraction rules, the template orchestrates specialized AI agents that scrape color palettes, typography, and spacing values directly from rendered web pages. This article examines the complete technical pipeline for how the template extracts design tokens from a target site.
The AI-Driven Extraction Architecture
The template does not hard-code design-token logic. Instead, it delegates extraction to AI agents that execute as part of the clone-website command. A foreman agent coordinates the workflow, checking for available Browser-MCP tools (such as Chrome MCP or Playwright MCP) to handle headless browser operations. This architecture allows the system to dynamically inspect any target site without requiring site-specific parsers.
The Six-Step Token Extraction Pipeline
The extraction process follows a strict runtime pipeline defined in the repository’s skill files and workflows.
Step 1: Launch the CLI and Initialize the Foreman
The process begins when the user runs npm run dev or invokes /clone-website <url>. This command boots the foreman agent, which coordinates the entire cloning operation. According to the repository’s documentation in README.md at Line 9, this entry point triggers the agentic workflow that manages downstream token extraction.
Step 2: Browser Automation and Page Rendering
The foreman agent first checks for a Browser-MCP capability (Chrome MCP, Playwright MCP, or equivalent). It then opens the target URL in a headless browser to fully render the DOM, including all JavaScript-generated styles. This step is documented in .github/skills/clone-website/SKILL.md at Lines 21‑23, which specifies the browser-automation prerequisites for the skill.
Step 3: CSS Inspection via Computed Styles
Once the page renders, the agent queries the computed style of every DOM node using window.getComputedStyle(element). As outlined in .windsurf/workflows/clone-website.md at Lines 56‑59, the agent collects:
- Color values (hex, rgb, hsl) for later conversion to Oklch tokens.
- Font families, sizes, and weights.
- Spacing values (margin, padding, gap).
- Border radius and shadow values.
This approach captures the final rendered styles rather than static CSS files, ensuring tokens reflect the actual visual appearance.
Step 4: Token Aggregation and Semantic Mapping
Raw values are de-duplicated and mapped to semantic names such as --color-primary and --font-base. The agent writes a structured JSON or YAML artifact representing the design-token set. This aggregation logic follows the guidelines in docs/research/INSPECTION_GUIDE.md, which defines the standard for naming and categorizing tokens.
Step 5: Tailwind v4 Configuration Generation
Using the aggregated token list, a utility builds a Tailwind CSS v4 configuration that defines the Oklch color palette, spacing scale, and font families. According to README.md at Lines 83‑84, this config becomes the global stylesheet for the cloned site, ensuring design consistency across generated components.
Step 6: Persistence and Component Integration
The generated token file (e.g., docs/research/<hostname>/design-tokens.json) and the Tailwind config are saved to the repository under docs/research/ and src/app/globals.css. As noted in .windsurf/workflows/clone-website.md at Lines 59‑61, subsequent builder agents read these files to produce pixel-perfect components that match the original site’s design system.
Implementation Deep Dive: Inspecting the DOM
The actual extraction relies on browser automation to inspect live DOM properties. The following TypeScript example illustrates the core logic implemented by the AI agents, using Playwright as a fallback when Chrome MCP is unavailable:
// 1️⃣ Open the page with Playwright
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
await page.goto(targetUrl);
// 2️⃣ Walk the DOM and collect computed styles
const tokens = await page.evaluate(() => {
const elems = Array.from(document.querySelectorAll('*'));
const map = new Set<string>();
elems.forEach(el => {
const cs = getComputedStyle(el);
// Extract key properties
map.add(cs.color);
map.add(cs.backgroundColor);
map.add(cs.fontFamily);
map.add(cs.fontSize);
map.add(cs.marginTop);
map.add(cs.padding);
// …additional properties as needed
});
return Array.from(map);
});
// 3️⃣ Transform raw values to Oklch tokens (simplified)
function toOklch(hex: string): string {
// Real implementation uses color-library (e.g., colorjs.io)
return hex; // placeholder for demo
}
const oklchTokens = tokens.map(v => toOklch(v));
// 4️⃣ Write the token file that the pipeline reads
import { writeFile } from 'node:fs/promises';
await writeFile(
`docs/research/${new URL(targetUrl).hostname}/design-tokens.json`,
JSON.stringify({ colors: oklchTokens }, null, 2)
);
Key Source Files and Their Responsibilities
The extraction pipeline is distributed across several configuration and documentation files:
README.md– High-level workflow description, mentions design-token extraction at Line 9 and Tailwind generation at Lines 83‑84.docs/research/INSPECTION_GUIDE.md– Detailed guide for inspection agents, including token-collection steps and semantic naming conventions..windsurf/workflows/clone-website.md– Procedural checklist outlining the prerequisite for creating "global CSS with the target site’s design tokens" (Lines 56‑61)..github/skills/clone-website/SKILL.md– AI skill definition that drives browser-automation and token extraction (Lines 21‑23).src/app/globals.css– Receives the generated Tailwind v4 utilities built from extracted tokens.scripts/sync-skills.mjs– Utility that syncs skill definitions, ensuring token-extraction logic remains current.
Summary
- The ai-website-cloner-template does not hard-code token extraction logic; it uses AI agents and browser automation.
- The foreman agent leverages Browser-MCP tools to render target sites and inspect computed styles via
window.getComputedStyle. - Extracted values (colors, fonts, spacing) are de-duplicated, mapped to semantic names, and stored as JSON/YAML.
- The pipeline generates a Tailwind v4 configuration using Oklch color tokens, persisted in
docs/research/andsrc/app/globals.css. - Source files including
.github/skills/clone-website/SKILL.mdand.windsurf/workflows/clone-website.mddefine the extraction protocol.
Frequently Asked Questions
What specific design properties does the template extract from a target site?
The template extracts color values (hex, rgb, hsl), font families and sizes, spacing measurements (margin, padding, gap), border radius, and shadow values. These properties are captured from the browser’s computed styles rather than static CSS files, ensuring the tokens reflect the final rendered appearance.
How does the template convert extracted colors to Oklch format?
After extracting raw color values via getComputedStyle, the AI agents pass these values through a color-conversion utility (typically using a library like colorjs.io). The conversion happens during the token aggregation phase before writing the design-tokens.json file, producing Oklch color values that feed into the Tailwind v4 configuration.
Where are the extracted design tokens stored within the repository?
The generated token files are saved to docs/research/<hostname>/design-tokens.json, where <hostname> corresponds to the target URL’s domain. Additionally, the generated Tailwind CSS configuration is written to src/app/globals.css, making the tokens available to all builder agents in the pipeline.
Can the extraction process function without Chrome MCP installed?
Yes. While the foreman agent prefers a Browser-MCP such as Chrome MCP for automation, the system falls back to Playwright or other available browser automation tools if the MCP is unavailable. This ensures the token extraction pipeline remains operational across different environment configurations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →