# How the Pre-Flight Checklist Ensures Complete Extraction Before Dispatching Builders

> Ensure complete extraction before dispatching builders with our pre-flight checklist. Validate browser automation, URL accessibility, and project integrity for verified extractions.

- Repository: [JCodesMore/ai-website-cloner-template](https://github.com/JCodesMore/ai-website-cloner-template)
- Tags: internals
- Published: 2026-07-07

---

**The pre-flight checklist validates browser automation, URL accessibility, project build integrity, and required directories before any builder agents receive component specifications, guaranteeing that every extraction is complete and verified.**

The `JCodesMore/ai-website-cloner-template` repository implements a rigorous pre-flight validation system within its `/clone-website` skill. This checkpoint ensures that the extraction phase has everything it needs before any builder agents are dispatched. By validating the environment, URLs, project build, and required directories up-front, the checklist guarantees that every subsequent builder receives a *complete, verified* specification rather than a partial or broken one.

## The Five-Step Pre-Flight Validation Process

The pre-flight checklist is defined in the skill file at lines 27-34 of [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md). It executes five mandatory validation steps that must all pass before the system proceeds to the reconnaissance phase.

### Browser Automation Detection

The checklist first detects available Model Context Protocol (MCP) tools for browser automation, including Chrome, Playwright, Browserbase, or Puppeteer. If no browser driver is found, it prompts the user to install one and aborts execution.

This matters because extraction relies on a live DOM to query `getComputedStyle()`, capture screenshots, and interact with pages. Without a functional browser MCP, the script would silently produce empty data, leaving builders with incomplete specifications.

### URL Parsing and Validation

Each supplied URL is normalized, syntax-checked, and verified for accessibility through the chosen browser tool. Invalid URLs immediately abort the run with a clear error message.

This early validation guarantees that the target site can actually be inspected. A bad URL would produce empty specs and break downstream builders, wasting compute cycles on unreachable targets.

### Project Build Verification

The checklist runs `npm run build` on the scaffolded Next.js 16 project to ensure the base code compiles cleanly. This verification happens before any extraction begins.

If the scaffold is broken, any spec that references components, utilities, or Tailwind classes would cause a build failure later. Verifying the build ensures that later imports—such as the `cn()` utility from [`src/lib/utils.ts`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/src/lib/utils.ts)—resolve correctly when builders execute their tasks.

### Output Directory Creation

The system ensures that `docs/research/`, `docs/research/components/`, `docs/design-references/`, and `scripts/` exist, including per-site subfolders for parallel runs.

Builder agents read specifications from `docs/research/components/`; missing folders would cause file-write errors and loss of audit artifacts. Pre-creating these directories means extraction scripts can write files atomically, eliminating race conditions that could truncate spec files.

### Parallel Execution Confirmation

When multiple URLs are provided, the checklist asks whether to run them in parallel or sequentially. This optional confirmation prevents resource overload.

Parallel extraction can overwhelm system resources or trigger rate limits. Confirming the execution mode allows users to stay in control and avoid incomplete extraction due to throttling or timeout errors.

## Why This Guarantees Complete Extraction

The pre-flight checklist ensures complete extraction through five defensive mechanisms:

1. **Environment sanity before DOM work** – Detecting a functional browser MCP prevents silent failures where the script would produce empty data.
2. **Early URL validation** – Invalid or unreachable URLs are caught before any network traffic, avoiding half-finished specs.
3. **Scaffold integrity** – Verifying the Next.js project builds guarantees that imports like `cn()` and shadcn primitives resolve correctly.
4. **Filesystem readiness** – Pre-creating output folders ensures extraction scripts can write files atomically without mid-process failures.
5. **Resource-aware parallelism** – By optionally toggling parallel execution, the checklist stops overload-driven timeouts that would leave sections un-extracted.

All checks are mandatory; the skill aborts with a clear error if any validation fails, forcing resolution before any builder agents are dispatched.

## Pre-Dispatch Verification Layer

After the initial pre-flight phase and subsequent extraction, a secondary "Pre-Dispatch Checklist" (lines 31-45 in [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md)) re-verifies that the spec is exhaustive.

This layer confirms that all CSS from `getComputedStyle()` is captured, the interaction model is documented, all states are captured, assets are listed, and responsive details are recorded. This double-layered guard ensures *both* the environment and the content are complete before any builder is launched.

## Implementing the Pre-Flight Logic

You can replicate the pre-flight validation in your own CI pipelines or local scripts. Below is a TypeScript implementation that mirrors the skill's logic:

```typescript
// preflight.ts – replicates the skill's validation logic
import { execSync } from "node:child_process";
import fetch from "node-fetch";

async function hasBrowserMCP(): Promise<boolean> {
  // Check for headless Chrome availability
  try {
    execSync("chromium --version", { stdio: "ignore" });
    return true;
  } catch {
    return false;
  }
}

function validateUrl(url: string): boolean {
  try {
    const u = new URL(url);
    return ["http:", "https:"].includes(u.protocol);
  } catch {
    return false;
  }
}

function verifyBuild(): boolean {
  try {
    execSync("npm run build", { stdio: "ignore" });
    return true;
  } catch {
    return false;
  }
}

function ensureDirs(): void {
  const dirs = [
    "docs/research",
    "docs/research/components",
    "docs/design-references",
    "scripts",
  ];
  for (const d of dirs) execSync(`mkdir -p ${d}`);
}

// ---- execution ----
(async () => {
  const urls = process.argv.slice(2);
  
  if (!(await hasBrowserMCP())) {
    console.error("❌ No browser MCP tool detected. Install Chrome, Playwright, etc.");
    process.exit(1);
  }

  for (const u of urls) {
    if (!validateUrl(u)) {
      console.error(`❌ Invalid URL: ${u}`);
      process.exit(1);
    }
    const res = await fetch(u, { method: "HEAD" });
    if (!res.ok) {
      console.error(`❌ URL not reachable (status ${res.status}): ${u}`);
      process.exit(1);
    }
  }

  if (!verifyBuild()) {
    console.error("❌ Project fails to build. Run `npm install` and fix errors first.");
    process.exit(1);
  }

  ensureDirs();
  console.log("✅ Pre-flight passed – extraction can safely begin.");
})();

```

Run the validation with:

```bash
npx ts-node preflight.ts https://example.com

```

This produces the same "all green" outcome the skill expects before launching extraction agents.

## Key Source Files in the Pipeline

| File | Role in the Pre-Flight & Extraction Pipeline |
|------|----------------------------------------------|
| [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md) | Defines the entire `/clone-website` workflow, including the pre-flight checklist (lines 27-34) and the pre-dispatch checklist (lines 31-45). |
| [`README.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/README.md) | High-level overview explaining the multi-phase pipeline that the pre-flight checklist gates. |
| [`AGENTS.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/AGENTS.md) | Central source of truth for agent-specific instructions; changes propagate to all agents through this file. |
| [`scripts/sync-agent-rules.sh`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/scripts/sync-agent-rules.sh) | Regenerates platform-specific instruction files after editing [`AGENTS.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/AGENTS.md). |
| `scripts/sync-skills.mjs` | Rebuilds the [`.claude/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.claude/skills/clone-website/SKILL.md) copy after the master skill file is edited. |
| [`src/lib/utils.ts`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/src/lib/utils.ts) | Provides the `cn()` utility used by builders; the build verification ensures this helper is present. |
| [`src/app/layout.tsx`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/src/app/layout.tsx) & [`src/app/globals.css`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/src/app/globals.css) | Global files updated during the Foundation Build phase; pre-flight guarantees they compile before spec generation. |

## Summary

- The pre-flight checklist in `JCodesMore/ai-website-cloner-template` validates browser automation, URLs, project builds, and directories before any extraction begins.
- Located at lines 27-34 of [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md), these checks are mandatory and abort on failure.
- A secondary pre-dispatch checklist (lines 31-45) verifies spec completeness before builders receive their tasks.
- The system prevents incomplete extraction by ensuring DOM tools are available, URLs are reachable, the Next.js scaffold compiles, and output directories exist.
- Builders only receive specifications after both validation layers pass, eliminating partial or broken component generation.

## Frequently Asked Questions

### What happens if no browser MCP tool is detected during pre-flight?

The skill aborts immediately with a clear error message prompting the user to install Chrome, Playwright, Browserbase, or Puppeteer. Without a browser driver, the system cannot query `getComputedStyle()` or capture screenshots, which would result in empty extraction data.

### Why does the checklist verify the project build before extraction?

The build verification ensures that the scaffolded Next.js 16 project compiles cleanly. This guarantees that utilities like the `cn()` function in [`src/lib/utils.ts`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/src/lib/utils.ts) and Tailwind configurations are valid before builders attempt to import them during component generation.

### What directories must exist before extraction can begin?

The checklist ensures `docs/research/`, `docs/research/components/`, `docs/design-references/`, and `scripts/` exist, along with per-site subfolders. Builder agents read specs from `docs/research/components/`; missing folders would cause file-write errors and loss of audit artifacts.

### How does the pre-dispatch checklist differ from the initial pre-flight validation?

While the pre-flight checklist validates the environment and project setup (lines 27-34), the pre-dispatch checklist (lines 31-45) verifies that the extracted specification itself is complete. It confirms all CSS, interaction models, states, assets, and responsive details are captured before any builder agent receives the component spec.