How the Asset Discovery Script Finds and Downloads Images and Videos in the AI Website Cloner
The asset discovery script executes inside the browser's page context to extract all image and video sources from the DOM, then downloads them to public/images/ and public/videos/ using Node.js streaming.
The JCodesMore/ai-website-cloner-template repository automates website replication by running a sophisticated asset discovery routine during the clone-website workflow. This process ensures every visual element—from standard image tags to CSS background images—is captured and stored locally for the generated Next.js application. The script operates in two phases: first querying the live DOM to catalog media URLs, then persisting those files to the filesystem with preserved directory hierarchies.
Execution Context and DOM Inspection
The asset discovery begins when scripts/download-assets.mjs executes inside the browser's page context opened by the AI agent. This approach allows direct access to the rendered DOM and computed CSS styles, ensuring no hidden or dynamically loaded assets are missed.
The script leverages the browser's native APIs rather than static HTML parsing, which enables it to capture images loaded via JavaScript and detect visual elements that only appear after user interaction or scroll events.
Discovering Media Elements in the Browser
The discovery phase systematically scans three categories of visual media elements using standard DOM selectors.
Extracting Image Sources
For standard image elements, the script runs document.querySelectorAll('img') to obtain a NodeList of every <img> tag on the page. For each element, it records:
- The
srcattribute containing the image URL - The
alttext for accessibility preservation - The
naturalWidthandnaturalHeightproperties to understand original dimensions
These values are collected into an array and de-duplicated before transmission to the Node.js download handler.
Capturing Video Sources
For video content, the script selects all <video> elements via document.querySelectorAll('video'). Unlike images, video elements often contain multiple <source> children for different formats, so the script flattens the results using:
const videos = [...document.querySelectorAll('video')].flatMap(v =>
[...v.querySelectorAll('source')].map(src => src.src)
);
This approach ensures all available video formats (MP4, WebM, OGG) are discovered even when the browser only displays one.
Detecting CSS Background Images
Many websites use CSS background images for decorative elements and hero sections. The script identifies these by iterating through all DOM elements and reading computed styles:
const backgrounds = [...document.querySelectorAll('*')]
.map(el => getComputedStyle(el).backgroundImage)
.filter(bg => bg && bg !== 'none')
.map(bg => bg.replace(/^url\(["']?/, '').replace(/["']?\)$/, ''));
All discovered URLs are normalized to absolute URLs and filtered to include only http and https schemes, excluding data URIs and local file references.
Downloading and Saving Assets
Once the browser-side collection completes, the URL arrays are passed to the Node.js side of scripts/download-assets.mjs for persistence.
Streaming Downloads to Local Storage
The script uses the built-in fetch API (with node-fetch as a fallback) to stream resources directly to disk, minimizing memory overhead for large video files:
import { pipeline } from 'stream/promises';
import fs from 'fs';
import path from 'path';
async function downloadAsset(url) {
const res = await fetch(url);
if (!res.ok) throw new Error(`Failed ${url}: ${res.statusText}`);
const isVideo = /\.(mp4|webm|ogg)$/i.test(url);
const folder = isVideo ? 'public/videos' : 'public/images';
const filename = path.basename(new URL(url).pathname);
const dest = path.join(folder, filename);
await pipeline(res.body, fs.createWriteStream(dest));
console.log(`✔︎ Saved ${url} → ${dest}`);
}
The destination directory is determined by file extension—images route to public/images/ while videos route to public/videos/. The script automatically handles HTTP redirects and follows them to completion.
Preserving Directory Structure
When an asset URL contains sub-paths (e.g., https://example.com/assets/icons/logo.svg), the script recreates the identical sub-directory hierarchy inside the appropriate public/ folder. This preservation ensures that relative path references in the original site's CSS and HTML remain valid after cloning.
If filename collisions occur during download, the script appends a numeric suffix to maintain distinct files without overwriting existing assets.
Integration with the Clone Workflow
After the download phase completes, the clone-website workflow (documented in .windsurf/workflows/clone-website.md) performs additional integration steps:
- SEO asset handling: Favicons and Open Graph images saved under
public/seo/are automatically referenced in the generatedlayout.tsxmetadata - Component updates: All
<Image>and<video>components in the cloned site are rewritten to point to the local/images/and/videos/paths rather than external URLs - Offline functionality: The generated Next.js application works entirely offline since all visual assets are stored locally
This integration ensures the final replica matches the original site pixel-perfectly while maintaining full functionality without external dependencies.
Summary
- Browser execution: The script runs inside the page context via
scripts/download-assets.mjsto access live DOM and computed styles - Comprehensive discovery: Captures
<img>tags,<video>sources, and CSSbackground-imageproperties usingquerySelectorAllandgetComputedStyle - Intelligent sorting: Routes downloads to
public/images/orpublic/videos/based on file extensions - Structure preservation: Recreates sub-directory hierarchies from original URLs to maintain relative path integrity
- Workflow integration: Automatically updates Next.js components and metadata to reference local assets
Frequently Asked Questions
How does the script handle images loaded by JavaScript after the initial page load?
The asset discovery script executes after the AI agent has allowed the page to fully render, including any lazy-loaded or dynamically injected content. Because it queries the live DOM using document.querySelectorAll('img') at runtime rather than parsing static HTML, it captures images injected by React, Vue, or vanilla JavaScript frameworks that appear after initial paint.
What video formats does the asset discovery script support?
The script recognizes MP4, WebM, and OGG files through regex pattern matching on the URL extension (/\.(mp4|webm|ogg)$/i). It extracts the src attribute from every <source> element nested within <video> tags, ensuring it captures all available format alternatives for cross-browser compatibility.
How does the script prevent duplicate downloads when the same image appears multiple times on a page?
Before transmitting URLs to the Node.js download handler, the script de-duplicates the collected arrays using JavaScript's Set data structure or explicit filtering. Additionally, the download logic checks for existing filenames in the destination directory and appends numeric suffixes when collisions occur, preventing overwrites while maintaining unique references.
Can the asset discovery script download assets from CSS files directly?
The script specifically targets inline CSS background images applied via the style attribute or classes by reading getComputedStyle(element).backgroundImage on every DOM element. While it does not parse external CSS files directly, it captures any background image that renders in the browser, including those defined in external stylesheets, because getComputedStyle resolves the final computed value regardless of the CSS source.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →