# How the Asset Discovery Script Finds and Downloads Images and Videos in the AI Website Cloner

> Learn how the ai website cloner asset discovery script finds and downloads images and videos directly from any webpage. Discover its efficient DOM extraction and Node.js streaming.

- Repository: [JCodesMore/ai-website-cloner-template](https://github.com/JCodesMore/ai-website-cloner-template)
- Tags: internals
- Published: 2026-07-10

---

**The asset discovery script executes inside the browser's page context to extract all image and video sources from the DOM, then downloads them to `public/images/` and `public/videos/` using Node.js streaming.**

The **JCodesMore/ai-website-cloner-template** repository automates website replication by running a sophisticated asset discovery routine during the `clone-website` workflow. This process ensures every visual element—from standard image tags to CSS background images—is captured and stored locally for the generated Next.js application. The script operates in two phases: first querying the live DOM to catalog media URLs, then persisting those files to the filesystem with preserved directory hierarchies.

## Execution Context and DOM Inspection

The asset discovery begins when `scripts/download-assets.mjs` executes **inside the browser's page context** opened by the AI agent. This approach allows direct access to the rendered DOM and computed CSS styles, ensuring no hidden or dynamically loaded assets are missed.

The script leverages the browser's native APIs rather than static HTML parsing, which enables it to capture images loaded via JavaScript and detect visual elements that only appear after user interaction or scroll events.

## Discovering Media Elements in the Browser

The discovery phase systematically scans three categories of visual media elements using standard DOM selectors.

### Extracting Image Sources

For **standard image elements**, the script runs `document.querySelectorAll('img')` to obtain a NodeList of every `<img>` tag on the page. For each element, it records:

- The `src` attribute containing the image URL
- The `alt` text for accessibility preservation
- The `naturalWidth` and `naturalHeight` properties to understand original dimensions

These values are collected into an array and de-duplicated before transmission to the Node.js download handler.

### Capturing Video Sources

For **video content**, the script selects all `<video>` elements via `document.querySelectorAll('video')`. Unlike images, video elements often contain multiple `<source>` children for different formats, so the script flattens the results using:

```javascript
const videos = [...document.querySelectorAll('video')].flatMap(v =>
  [...v.querySelectorAll('source')].map(src => src.src)
);

```

This approach ensures all available video formats (MP4, WebM, OGG) are discovered even when the browser only displays one.

### Detecting CSS Background Images

Many websites use **CSS background images** for decorative elements and hero sections. The script identifies these by iterating through all DOM elements and reading computed styles:

```javascript
const backgrounds = [...document.querySelectorAll('*')]
  .map(el => getComputedStyle(el).backgroundImage)
  .filter(bg => bg && bg !== 'none')
  .map(bg => bg.replace(/^url\(["']?/, '').replace(/["']?\)$/, ''));

```

All discovered URLs are normalized to absolute URLs and filtered to include only `http` and `https` schemes, excluding data URIs and local file references.

## Downloading and Saving Assets

Once the browser-side collection completes, the URL arrays are passed to the Node.js side of `scripts/download-assets.mjs` for persistence.

### Streaming Downloads to Local Storage

The script uses the built-in `fetch` API (with `node-fetch` as a fallback) to stream resources directly to disk, minimizing memory overhead for large video files:

```javascript
import { pipeline } from 'stream/promises';
import fs from 'fs';
import path from 'path';

async function downloadAsset(url) {
  const res = await fetch(url);
  if (!res.ok) throw new Error(`Failed ${url}: ${res.statusText}`);

  const isVideo = /\.(mp4|webm|ogg)$/i.test(url);
  const folder = isVideo ? 'public/videos' : 'public/images';
  const filename = path.basename(new URL(url).pathname);
  const dest = path.join(folder, filename);

  await pipeline(res.body, fs.createWriteStream(dest));
  console.log(`✔︎ Saved ${url} → ${dest}`);
}

```

The destination directory is determined by file extension—images route to `public/images/` while videos route to `public/videos/`. The script automatically handles HTTP redirects and follows them to completion.

### Preserving Directory Structure

When an asset URL contains sub-paths (e.g., `https://example.com/assets/icons/logo.svg`), the script recreates the identical sub-directory hierarchy inside the appropriate `public/` folder. This preservation ensures that relative path references in the original site's CSS and HTML remain valid after cloning.

If filename collisions occur during download, the script appends a numeric suffix to maintain distinct files without overwriting existing assets.

## Integration with the Clone Workflow

After the download phase completes, the `clone-website` workflow (documented in [`.windsurf/workflows/clone-website.md`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/.windsurf/workflows/clone-website.md)) performs additional integration steps:

- **SEO asset handling**: Favicons and Open Graph images saved under `public/seo/` are automatically referenced in the generated [`layout.tsx`](https://github.com/JCodesMore/ai-website-cloner-template/blob/main/layout.tsx) metadata
- **Component updates**: All `<Image>` and `<video>` components in the cloned site are rewritten to point to the local `/images/` and `/videos/` paths rather than external URLs
- **Offline functionality**: The generated Next.js application works entirely offline since all visual assets are stored locally

This integration ensures the final replica matches the original site pixel-perfectly while maintaining full functionality without external dependencies.

## Summary

- **Browser execution**: The script runs inside the page context via `scripts/download-assets.mjs` to access live DOM and computed styles
- **Comprehensive discovery**: Captures `<img>` tags, `<video>` sources, and CSS `background-image` properties using `querySelectorAll` and `getComputedStyle`
- **Intelligent sorting**: Routes downloads to `public/images/` or `public/videos/` based on file extensions
- **Structure preservation**: Recreates sub-directory hierarchies from original URLs to maintain relative path integrity
- **Workflow integration**: Automatically updates Next.js components and metadata to reference local assets

## Frequently Asked Questions

### How does the script handle images loaded by JavaScript after the initial page load?

The asset discovery script executes after the AI agent has allowed the page to fully render, including any lazy-loaded or dynamically injected content. Because it queries the live DOM using `document.querySelectorAll('img')` at runtime rather than parsing static HTML, it captures images injected by React, Vue, or vanilla JavaScript frameworks that appear after initial paint.

### What video formats does the asset discovery script support?

The script recognizes MP4, WebM, and OGG files through regex pattern matching on the URL extension (`/\.(mp4|webm|ogg)$/i`). It extracts the `src` attribute from every `<source>` element nested within `<video>` tags, ensuring it captures all available format alternatives for cross-browser compatibility.

### How does the script prevent duplicate downloads when the same image appears multiple times on a page?

Before transmitting URLs to the Node.js download handler, the script de-duplicates the collected arrays using JavaScript's Set data structure or explicit filtering. Additionally, the download logic checks for existing filenames in the destination directory and appends numeric suffixes when collisions occur, preventing overwrites while maintaining unique references.

### Can the asset discovery script download assets from CSS files directly?

The script specifically targets inline CSS background images applied via the `style` attribute or classes by reading `getComputedStyle(element).backgroundImage` on every DOM element. While it does not parse external CSS files directly, it captures any background image that renders in the browser, including those defined in external stylesheets, because `getComputedStyle` resolves the final computed value regardless of the CSS source.