How the Asset Download Script in the AI Website Cloner Template Finds and Downloads All Media
The asset download script in the JCodesMore/ai-website-cloner-template operates as an integrated skill within the /clone-website pipeline, using browser automation to scan DOM elements and CSS properties, normalize URLs, and stream media files into categorized public/ directories.
The AI Website Cloner Template handles media retrieval through an intelligent, agent-driven process rather than a standalone shell script. According to the repository's architecture, the download logic lives inside the /clone-website skill that executes during the Foundation phase of the cloning pipeline. This approach ensures that every image, video, and background asset referenced by the target site is faithfully extracted and stored in the generated Next.js project structure.
How the Asset Download Script Works
The media collection system is embedded within the skill definition found in .github/skills/clone-website/SKILL.md (generated from AGENTS.md). When the AI agent invokes this skill, it orchestrates a multi-step pipeline that combines browser automation with Node.js streaming capabilities. Unlike traditional wget or curl scripts, this implementation leverages the same browser context used for rendering, ensuring it captures dynamically loaded assets and computed CSS backgrounds that static parsers often miss.
Step-by-Step Media Discovery Process
DOM Scanning and CSS Extraction
The skill initiates by controlling the target page through a browser automation tool—typically Chrome MCP or Playwright MCP. It executes two parallel collection strategies:
- Element attribute extraction – The script queries all
<img>,<video>,<source>, and<audio>tags, extracting theirsrcandhrefattributes. - Computed style inspection – It iterates through all DOM elements, reading computed styles to capture
background-image: url(...)declarations.
This dual approach ensures coverage of both semantic media elements and decorative CSS-loaded assets.
URL Normalization and Deduplication
Once extracted, raw URLs undergo normalization to handle relative paths, protocol-relative links, and data-URL fallbacks. The collectMedia() function converts all references to absolute URLs before passing them to a Set constructor to eliminate duplicates. This deduplication step prevents redundant network requests and ensures each unique asset is downloaded exactly once, even if referenced multiple times across the site.
Download and Storage Mechanism
HTTP Request Handling
For each unique URL, the script performs an HTTP GET request using Node.js built-in https/http modules or a lightweight fetch wrapper. The response stream pipes directly to the filesystem using streamPipeline(), minimizing memory overhead during large file transfers. Error handling skips broken links gracefully, allowing the cloning process to continue even if individual assets are unreachable.
Directory Structure and File Naming
The script categorizes downloads based on file extensions:
- Images →
public/images/(e.g.,public/images/logo.png) - Videos →
public/videos/(e.g.,public/videos/hero.mp4)
File names derive from the URL's pathname via path.basename(). If naming conflicts arise, the script appends numeric suffixes to prevent overwrites. This structure aligns with the project layout documented in README.md, where these directories serve as the canonical storage for cloned assets.
Asset Manifest Generation
After completing downloads, the skill writes a JSON manifest to public/assets-manifest.json. This file maps original source URLs to their new local paths, enabling the subsequent component generation phase to reference correct relative paths in the React code. The manifest acts as the single source of truth for asset remapping, ensuring that generated components point to public/images/ and public/videos/ rather than external URLs.
Code Implementation Example
The following TypeScript pseudocode illustrates the core logic implemented in the skill's execution flow:
// Pseudocode extracted from the skill's internal implementation
async function collectMedia(page: Page): Promise<string[]> {
// 1️⃣ Grab all <img>, <video>, <source>, <audio> elements
const srcs = await page.$$eval(
'img, video, source, audio',
els => els.map(e => (e as HTMLMediaElement).src)
);
// 2️⃣ Pull background images from computed styles
const cssUrls = await page.evaluate(() => {
const urls = new Set<string>();
for (const el of document.querySelectorAll<HTMLElement>('*')) {
const bg = getComputedStyle(el).backgroundImage;
const match = bg.match(/url\(["']?([^"')]+)["']?\)/);
if (match) urls.add(match[1]);
}
return Array.from(urls);
});
// Combine and dedupe
return Array.from(new Set([...srcs, ...cssUrls])).filter(Boolean);
}
async function downloadAll(mediaUrls: string[]) {
for (const url of mediaUrls) {
const resp = await fetch(url);
if (!resp.ok) continue; // skip broken links
const urlObj = new URL(url);
const ext = path.extname(urlObj.pathname);
const destFolder = ext.match(/\.(mp4|webm|ogg)$/i)
? 'public/videos'
: 'public/images';
const destPath = path.join(destFolder, path.basename(urlObj.pathname));
await streamPipeline(resp.body, fs.createWriteStream(destPath));
}
}
The actual implementation is generated on-the-fly by the AI agent based on the skill definition, but follows this exact logical flow documented in docs/research/INSPECTION_GUIDE.md.
Summary
- The asset download script is integrated into the
/clone-websiteskill, not a standalone file. - It scans both DOM elements (
img,video,audio) and computed CSS (background-image) to discover all media. - URLs are normalized to absolute paths and deduplicated before downloading.
- Assets stream to
public/images/andpublic/videos/with automatic conflict resolution. - A
public/assets-manifest.jsonfile maps original URLs to local paths for component generation.
Frequently Asked Questions
Where is the asset download script located in the repository?
The download logic resides in the .github/skills/clone-website/SKILL.md file, which is generated from AGENTS.md. Unlike traditional templates that include a separate download-media.sh script, this template embeds the functionality within an AI skill that executes during the Foundation phase of the cloning pipeline.
How does the script handle CSS background images?
The script executes page.evaluate() to iterate through all DOM elements, calling getComputedStyle() on each to extract background-image URLs. It uses regex matching to parse the url() declarations, ensuring decorative images loaded via CSS are captured alongside standard <img> tags.
What happens if the target site has broken media links?
The downloadAll() function checks the response status for each HTTP request. If resp.ok returns false, the script skips that specific URL and continues processing the remaining assets. This fault-tolerant approach ensures that one missing image does not halt the entire cloning operation.
Does the script support video and audio files?
Yes. The asset download script explicitly targets <video>, <source>, and <audio> tags during the DOM scan. It routes files with extensions matching /\.(mp4|webm|ogg)$/i to the public/videos/ directory, while all other media types default to public/images/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →