The Five Phases of the Clone-Website Pipeline Explained
The five phases of the clone-website pipeline follow a structured workflow—Reconnaissance, Foundation, Component Specs, Parallel Build, and Assembly & QA—that guides AI agents from initial site inspection to a pixel-perfect Next.js deployment using isolated Git worktrees and incremental verification.
The JCodesMore/ai-website-cloner-template repository implements the five phases of the clone-website pipeline through the /clone-website skill. By breaking the cloning process into distinct, verifiable stages, the system ensures reproducible builds, isolates component failures, and delivers maintainable high-fidelity replicas of any target website.
Pipeline Architecture Overview
The complete workflow progresses through five sequential phases, with each stage producing auditable artifacts that feed into the next:
flowchart LR
A[Reconnaissance] --> B[Foundation]
B --> C[Component Specs]
C --> D[Parallel Build]
D --> E[Assembly & QA]
According to the source documentation in README.md (lines 90-95), this separation ensures that extraction, specification, and construction remain independent concerns, minimizing guesswork while maximizing parallelism.
Phase 1: Reconnaissance
The Reconnaissance phase performs comprehensive extraction of the target site's visual and behavioral DNA. This phase captures full-page screenshots, extracts global design tokens—including fonts, colors, and favicons—and executes an exhaustive interaction sweep.
The interaction sweep maps every dynamic element by simulating scroll, click, hover, and responsive behaviors across device breakpoints. Extracted data populates docs/design-references/ and docs/research/BEHAVIORS.md, creating a complete audit trail of the original site's functionality.
Phase 2: Foundation
Following extraction, the Foundation phase updates the project's global assets to match the target design system. The pipeline modifies src/app/layout.tsx to inject the extracted font families, ensuring typographic consistency across the application.
Simultaneously, the system updates src/app/globals.css with the extracted color palette and global CSS utilities. Helper scripts like scripts/download-assets.mjs pull down site-wide assets—including images, videos, and icons—into the public/ folder, while TypeScript interfaces are generated to support type-safe component development.
Phase 3: Component Specs
The Component Specs phase generates detailed specification files within docs/research/components/, serving as architectural contracts for downstream builders. Each specification captures exact computed CSS values, interaction models, multi-state content variations, and precise asset paths.
These specification files eliminate ambiguity by documenting precise implementation requirements for every UI element. By separating specification from construction, the pipeline enables parallel development while maintaining design fidelity to the original screenshots.
Phase 4: Parallel Build
The Parallel Build phase leverages Git worktrees to spawn isolated builder agents—one per component or section—enabling concurrent React/TSX implementation. As defined in /.github/skills/clone-website/SKILL.md, each builder works in isolation while referencing the component specs created in Phase 3.
This parallelization strategy dramatically reduces build time while preventing merge conflicts. Because each component evolves in its own worktree, builders can experiment with implementation details without affecting the main codebase or other concurrent development streams.
Phase 5: Assembly & QA
The final Assembly & QA phase merges all worktrees back into the main codebase and wires components together into a cohesive application. The pipeline executes a visual diff comparison against the original Phase 1 screenshots to verify pixel-perfect reproduction.
The system runs npm run build to ensure production readiness and validates that all assets resolve correctly. This incremental verification guarantees that the final clone matches the source site behaviorally and visually before deployment.
Executing the Pipeline
To trigger the five-phase pipeline, start the AI agent and invoke the clone skill:
# Launch Claude Code with browser integration
claude --chrome
# Initiate the five-phase pipeline
/clone-website https://example.com https://another-site.org
You can inspect intermediate artifacts after each phase completes:
# Phase 1 artifacts: screenshots and behavior mappings
ls docs/design-references/
cat docs/research/BEHAVIORS.md
# Phase 2 artifacts: global styles and assets
cat src/app/layout.tsx
cat src/app/globals.css
ls public/
# Phase 3 artifacts: component specifications
ls docs/research/components/
# Phase 4 artifacts: active builder worktrees
git worktree list
Summary
- Reconnaissance extracts visual tokens and behavioral mappings through automated interaction sweeps, populating
docs/research/. - Foundation applies global assets to
src/app/layout.tsx,src/app/globals.css, and thepublic/directory. - Component Specs generates detailed build contracts in
docs/research/components/that capture exact CSS and interaction requirements. - Parallel Build utilizes Git worktrees for isolated, concurrent component construction without merge conflicts.
- Assembly & QA merges worktrees, verifies pixel-perfect accuracy via visual diff, and validates the production build with
npm run build.
Frequently Asked Questions
What triggers the transition between phases in the clone-website pipeline?
Each phase completes only after specific verification gates defined in /.github/skills/clone-website/SKILL.md are satisfied. The Reconnaissance phase requires successful screenshot capture and token extraction, while the Parallel Build phase waits for all component specifications to be generated before spawning builder agents.
Why does the pipeline use Git worktrees instead of traditional branches?
Git worktrees provide physical filesystem isolation that prevents cross-contamination between concurrent component builds. Unlike branches, worktrees allow multiple builders to modify shared dependencies simultaneously without merge conflicts, and they enable the Assembly phase to merge completed components through simple directory operations rather than complex Git merges.
How does the Assembly & QA phase ensure pixel-perfect accuracy?
The final phase performs a visual diff comparison between the built application and the original screenshots captured during Phase 1. It also executes npm run build to verify that all imports, assets, and TypeScript definitions resolve correctly, ensuring the clone matches the source site both visually and functionally.
Where are the component specifications stored during the pipeline?
Phase 3 generates specification files in the docs/research/components/ directory. These files contain computed CSS values, interaction models, and asset paths, serving as the single source of truth that Phase 4 builders reference when constructing React components in their isolated worktrees.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →