# How the Asset Download Script in the AI Website Cloner Template Finds and Downloads All Media

> Discover how the AI website cloner template's asset download script finds and downloads all website media. Learn about its DOM scanning, URL normalization, and file streaming capabilities.

- Repository: [JCodesMore/ai-website-cloner-template](https://github.com/JCodesMore/ai-website-cloner-template)
- Tags: how-to-guide
- Published: 2026-07-07

---

**The asset download script in the JCodesMore/ai-website-cloner-template operates as an integrated skill within the `/clone-website` pipeline, using browser automation to scan DOM elements and CSS properties, normalize URLs, and stream media files into categorized `public/` directories.**

The **AI Website Cloner Template** handles media retrieval through an intelligent, agent-driven process rather than a standalone shell script. According to the repository's architecture, the download logic lives inside the **`/clone-website` skill** that executes during the *Foundation* phase of the cloning pipeline. This approach ensures that every image, video, and background asset referenced by the target site is faithfully extracted and stored in the generated Next.js project structure.

## How the Asset Download Script Works

The media collection system is embedded within the skill definition found in [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md) (generated from [`AGENTS.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/AGENTS.md)). When the AI agent invokes this skill, it orchestrates a multi-step pipeline that combines browser automation with Node.js streaming capabilities. Unlike traditional wget or curl scripts, this implementation leverages the same browser context used for rendering, ensuring it captures dynamically loaded assets and computed CSS backgrounds that static parsers often miss.

## Step-by-Step Media Discovery Process

### DOM Scanning and CSS Extraction

The skill initiates by controlling the target page through a browser automation tool—typically Chrome MCP or Playwright MCP. It executes two parallel collection strategies:

1. **Element attribute extraction** – The script queries all `<img>`, `<video>`, `<source>`, and `<audio>` tags, extracting their `src` and `href` attributes.
2. **Computed style inspection** – It iterates through all DOM elements, reading computed styles to capture `background-image: url(...)` declarations.

This dual approach ensures coverage of both semantic media elements and decorative CSS-loaded assets.

### URL Normalization and Deduplication

Once extracted, raw URLs undergo normalization to handle relative paths, protocol-relative links, and data-URL fallbacks. The `collectMedia()` function converts all references to absolute URLs before passing them to a `Set` constructor to eliminate duplicates. This deduplication step prevents redundant network requests and ensures each unique asset is downloaded exactly once, even if referenced multiple times across the site.

## Download and Storage Mechanism

### HTTP Request Handling

For each unique URL, the script performs an HTTP GET request using Node.js built-in `https`/`http` modules or a lightweight fetch wrapper. The response stream pipes directly to the filesystem using `streamPipeline()`, minimizing memory overhead during large file transfers. Error handling skips broken links gracefully, allowing the cloning process to continue even if individual assets are unreachable.

### Directory Structure and File Naming

The script categorizes downloads based on file extensions:

- **Images** → `public/images/` (e.g., `public/images/logo.png`)
- **Videos** → `public/videos/` (e.g., `public/videos/hero.mp4`)

File names derive from the URL's pathname via `path.basename()`. If naming conflicts arise, the script appends numeric suffixes to prevent overwrites. This structure aligns with the project layout documented in [`README.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/README.md), where these directories serve as the canonical storage for cloned assets.

## Asset Manifest Generation

After completing downloads, the skill writes a JSON manifest to [`public/assets-manifest.json`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/public/assets-manifest.json). This file maps original source URLs to their new local paths, enabling the subsequent component generation phase to reference correct relative paths in the React code. The manifest acts as the single source of truth for asset remapping, ensuring that generated components point to `public/images/` and `public/videos/` rather than external URLs.

## Code Implementation Example

The following TypeScript pseudocode illustrates the core logic implemented in the skill's execution flow:

```typescript
// Pseudocode extracted from the skill's internal implementation

async function collectMedia(page: Page): Promise<string[]> {
  // 1️⃣ Grab all <img>, <video>, <source>, <audio> elements
  const srcs = await page.$$eval(
    'img, video, source, audio',
    els => els.map(e => (e as HTMLMediaElement).src)
  );

  // 2️⃣ Pull background images from computed styles
  const cssUrls = await page.evaluate(() => {
    const urls = new Set<string>();
    for (const el of document.querySelectorAll<HTMLElement>('*')) {
      const bg = getComputedStyle(el).backgroundImage;
      const match = bg.match(/url\(["']?([^"')]+)["']?\)/);
      if (match) urls.add(match[1]);
    }
    return Array.from(urls);
  });

  // Combine and dedupe
  return Array.from(new Set([...srcs, ...cssUrls])).filter(Boolean);
}

async function downloadAll(mediaUrls: string[]) {
  for (const url of mediaUrls) {
    const resp = await fetch(url);
    if (!resp.ok) continue;                     // skip broken links

    const urlObj = new URL(url);
    const ext = path.extname(urlObj.pathname);
    const destFolder = ext.match(/\.(mp4|webm|ogg)$/i)
      ? 'public/videos'
      : 'public/images';
    const destPath = path.join(destFolder, path.basename(urlObj.pathname));

    await streamPipeline(resp.body, fs.createWriteStream(destPath));
  }
}

```

The actual implementation is generated on-the-fly by the AI agent based on the skill definition, but follows this exact logical flow documented in [`docs/research/INSPECTION_GUIDE.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/docs/research/INSPECTION_GUIDE.md).

## Summary

- The **asset download script** is integrated into the `/clone-website` skill, not a standalone file.
- It scans both **DOM elements** (`img`, `video`, `audio`) and **computed CSS** (`background-image`) to discover all media.
- URLs are **normalized to absolute paths** and **deduplicated** before downloading.
- Assets stream to **`public/images/`** and **`public/videos/`** with automatic conflict resolution.
- A **[`public/assets-manifest.json`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/public/assets-manifest.json)** file maps original URLs to local paths for component generation.

## Frequently Asked Questions

### Where is the asset download script located in the repository?

The download logic resides in the [`.github/skills/clone-website/SKILL.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.github/skills/clone-website/SKILL.md) file, which is generated from [`AGENTS.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/AGENTS.md). Unlike traditional templates that include a separate [`download-media.sh`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/download-media.sh) script, this template embeds the functionality within an AI skill that executes during the *Foundation* phase of the cloning pipeline.

### How does the script handle CSS background images?

The script executes `page.evaluate()` to iterate through all DOM elements, calling `getComputedStyle()` on each to extract `background-image` URLs. It uses regex matching to parse the `url()` declarations, ensuring decorative images loaded via CSS are captured alongside standard `<img>` tags.

### What happens if the target site has broken media links?

The `downloadAll()` function checks the response status for each HTTP request. If `resp.ok` returns false, the script skips that specific URL and continues processing the remaining assets. This fault-tolerant approach ensures that one missing image does not halt the entire cloning operation.

### Does the script support video and audio files?

Yes. The asset download script explicitly targets `<video>`, `<source>`, and `<audio>` tags during the DOM scan. It routes files with extensions matching `/\.(mp4|webm|ogg)$/i` to the `public/videos/` directory, while all other media types default to `public/images/`.