How the Asset Discovery Script Downloads Layered Images in ai-website-cloner-template
TLDR: The asset discovery script in scripts/download-assets.mjs treats layered images as collections of individual assets, detecting each <img> tag and CSS background-image URL separately, queuing them with order preservation, and downloading them concurrently with a default concurrency of four to the public/ folder while maintaining the original path hierarchy.
The JCodesMore/ai-website-cloner-template repository automates website cloning by capturing every binary resource required for visual fidelity. At the core of this process is the asset discovery script that handles how layered images are downloaded and structured locally. This Node.js implementation ensures that composite visual elements—built from multiple stacked images—are preserved as distinct files that can be reassembled to match the original layout exactly.
Detecting Layered Image Compositions
The script identifies layered images by analyzing both HTML structure and computed styles, treating each visual constituent as a separate download candidate rather than a single opaque asset.
DOM Inspection Strategy
According to the workflow documentation in .windsurf/workflows/clone-website.md (lines 198-202), the script is instructed to "Include parent info to detect layered compositions." This guidance directs the discovery logic to examine container elements and their descendants holistically. The implementation walks the DOM tree recursively, accumulating layer information as it traverses parent-child relationships to understand how individual images stack together.
CSS Background Image Parsing
Beyond standard <img> tags, the script inspects computed CSS styles for background-image declarations. When encountering comma-separated URLs within a single background property, it extracts each URL as a distinct layer. This ensures that complex CSS sprites and multi-layered backgrounds are fully captured.
The Three-Stage Download Process
The asset discovery script follows a deterministic pipeline to ensure complete and ordered retrieval of layered visual assets.
Stage 1: Extracting Visual Layers from the DOM
The script implements a collectLayers function that traverses the DOM hierarchy. For each element, it records:
- All descendant
<img>tags viaquerySelectorAll('img'), capturing theirsrcattributes. - Any CSS
background-imagevalues usinggetComputedStyle, parsing multiple comma-separated URLs with regex matching.
This data collection preserves the depth and nesting information required to reconstruct the exact visual stacking order later.
Stage 2: Queueing Assets with Order Preservation
Each discovered URL—whether from an image source or background declaration—is pushed onto a download queue as a separate entry. The queue maintains the DOM traversal order, guaranteeing that the visual z-index and layering relationships can be replicated during the build phase. This approach treats a "layered image" not as a single file, but as an ordered collection of assets.
Stage 3: Parallel Download and File System Organization
The script processes the queue using a concurrency pool limited to four simultaneous fetches to avoid overwhelming the target host. For each asset:
- It uses the native
fetchAPI (with anhttps.getfallback) to stream binary data. - It writes files to the
public/directory while mirroring the original URL path hierarchy (e.g.,public/images/foo/bar.png). - It emits log entries upon completion to verify that all layers of a composite image were retrieved.
Implementation in scripts/download-assets.mjs
The core logic resides in scripts/download-assets.mjs, which implements the layer detection and download pipeline.
The layer collection logic extracts both inline images and background layers:
// walk each element; record img and background layers
function collectLayers(node, parentInfo = []) {
const layers = [...parentInfo];
// 1️⃣ <img> tags
node.querySelectorAll('img').forEach(img => {
const src = img.getAttribute('src');
if (src) layers.push({ url: src, type: 'img', depth: layers.length });
});
// 2️⃣ CSS background-image (may contain multiple URLs)
const style = getComputedStyle(node);
const bg = style.getPropertyValue('background-image');
if (bg && bg !== 'none') {
const urls = bg.match(/url\(["']?(.*?)["']?\)/g).map(u => u.slice(5, -2));
urls.forEach(url => layers.push({ url, type: 'bg', depth: layers.length }));
}
// recurse into children, passing the accumulated layer list
node.children.forEach(child => collectLayers(child, layers));
return layers;
}
The parallel download implementation uses a queue with concurrency control:
const concurrency = 4;
const queue = new PQueue({ concurrency });
for (const asset of assets) {
queue.add(async () => {
const res = await fetch(asset.url);
if (!res.ok) throw new Error(`Failed ${asset.url}`);
const dest = path.join('public', new URL(asset.url).pathname);
await fs.promises.mkdir(path.dirname(dest), { recursive: true });
const file = fs.createWriteStream(dest);
await new Promise((resolve, reject) => {
res.body.pipe(file).on('finish', resolve).on('error', reject);
});
console.log(`✔️ Layer downloaded: ${asset.url}`);
});
}
await queue.onIdle();
These implementations demonstrate how the script enumerates every visual layer, queues them independently, and downloads them concurrently while preserving the order necessary for accurate reconstruction.
Summary
- Layered images are decomposed: The asset discovery script treats composite visuals as individual
<img>andbackground-imageassets rather than single files. - Order preservation matters: The download queue maintains DOM traversal order to ensure accurate z-index and stacking reconstruction.
- Concurrency is capped: Downloads are limited to four simultaneous fetches to prevent server overload while maintaining efficient throughput.
- Path hierarchy is mirrored: Assets are stored in
public/following their original URL structure to maintain reference integrity. - Workflow guidance: The detection strategy is explicitly documented in
.windsurf/workflows/clone-website.mdto ensure AI agents capture all visual constituents.
Frequently Asked Questions
How does the script handle multiple background images in a single CSS property?
The script uses regex matching to parse comma-separated URLs within the background-image CSS value. Each URL extracted from the declaration is treated as a separate layer entry in the download queue, ensuring that complex multi-layered backgrounds are fully captured as individual assets.
What happens if an image download fails during the parallel fetch?
If a fetch request returns a non-OK status, the script throws an error with the specific URL that failed. This halts the queue processing for that asset, preventing incomplete data from being written to the file system and ensuring that failed layers are logged for debugging.
Why does the script preserve the original URL path hierarchy in the public folder?
Mirroring the original path structure (e.g., public/images/foo/bar.png) ensures that relative references in the cloned HTML and CSS remain valid without requiring complex URL rewriting. This approach maintains the integrity of the site's link structure and simplifies the reconstruction process.
Where is the layered image detection logic documented for AI agents?
The detection requirements are specified in .windsurf/workflows/clone-website.md at lines 198-202, which explicitly instructs agents to "Include parent info to detect layered compositions." This documentation ensures that the asset discovery process consistently identifies and captures all constituent layers of complex visual elements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →