Download Workflow from Request to ZIP Delivery: Technical Architecture of the Website Downloader

The download workflow in AhmadIbrahiim/Website-downloader implements a Socket.io-driven pipeline where client requests trigger a wget child process to mirror websites, with stderr progress streamed back to the client in real-time before the archiver module compresses the result into a ZIP file stored in public/sites/ for HTTP retrieval.

The repository provides a Node.js service that converts live websites into offline ZIP archives. Understanding the download workflow from request to ZIP delivery reveals how WebSocket events, child processes, and file system operations coordinate to deliver complete website mirrors with real-time progress updates.

Client Initiates Download via WebSocket

The workflow begins when the browser opens a Socket.io connection and emits a request event containing a unique token (to correlate responses) and the target website URL.

// Client-side implementation
const socket = io();
const token = Date.now();

socket.emit('request', {
  token,
  website: 'https://example.com'
});

socket.on(token, data => {
  console.log('Progress:', data.progress);
  if (data.file) {
    window.location = `/sites/${data.file}.zip`;
  }
});

The client listens on a channel named after the token to receive real-time progress updates and the final download URL.

Server Receives and Routes the Request

File: socket/socket.js

The server-side Socket.io listener captures the request and immediately forwards it to the download worker:

socket.on('request', function (data) {
  console.log("Request connection received %s", data.token);
  wget(io, data);  // Delegates to wget/index.js
});

The global io instance (Socket.io server) and request data are passed to the wget() function, establishing the connection between the real-time messaging layer and the file system operations.

wget Mirrors the Website with Real-Time Progress Streaming

File: wget/index.js

The download workflow spawns a child process using Node.js exec() with specific wget flags designed for complete website mirroring:

const child = exec(`wget -mkEpnp --no-if-modified-since ${data.website}`);

The -mkEpnp flags configure wget to create a recursive mirror, convert links for offline use, preserve file extensions, and avoid ascending to parent directories during the crawl.

Progress Transmission

All stderr output streams back to the client via Socket.io events, enabling live progress bars in the UI:

child.stderr.on("data", (response) => {
  const txt = response.toString();
  // Extract domain name from "Resolving..." output
  if (!website) {
    const m = txt.match(/Resolving\s+([^\s]+)\s+\(/);
    if (m) website = m[1];
  }
  io.emit(data.token, { progress: txt });
});

Folder Detection and Handoff

When wget completes, the close event handler determines the downloaded folder name and triggers the archiving phase:

child.stderr.on('close', () => {
  const websiteFolder = website || getWebsiteFolderName(data.website);
  io.emit(data.token, {progress: "Converting"});
  archiver(websiteFolder, io, data);  // Transition to zipping stage
});

The getWebsiteFolderName() utility extracts the hostname (or host:port) from the original URL if wget's output parsing failed to identify the domain.

Archiving the Downloaded Directory into ZIP

File: archiver/index.js

The archiver module creates a compressed ZIP file with maximum compression settings:

var output = fs.createWriteStream("./public/sites/" + file + '.zip');
var archive = archiver('zip', { zlib: { level: 9 } });

output.on('close', () => {
  console.log(archive.pointer() + ' total bytes');
  io.emit(data.token, { progress: "Completed", file });
});

archive.pipe(output);
archive.directory('./' + file, false);  // Include entire site folder
archive.finalize();

The level: 9 zlib setting applies maximum compression. The archive pipes to ./public/sites/<file>.zip, making it accessible via HTTP because app.js serves the public directory statically using express.static().

Error Handling and Cleanup

The download workflow includes mechanisms for handling client disconnections and partial failures. In socket/socket.js, disconnect events attempt to terminate running wget processes and abort archiver operations. The wget/index.js module implements removePartiallyDownloadedFiles to clean up incomplete downloads when processes terminate unexpectedly.

Both wget and archiver forward non-critical warnings while throwing on fatal errors, preventing silent failures in the pipeline.

Summary

  • WebSocket initiation: Clients emit request events with unique tokens and target URLs via Socket.io.
  • Server routing: socket/socket.js receives requests and delegates to wget/index.js.
  • Mirroring process: wget spawns with -mkEpnp flags to create offline mirrors while streaming stderr progress back to the client.
  • Folder detection: The system parses wget output or extracts hostnames to identify the downloaded directory.
  • ZIP creation: archiver/index.js compresses the folder into public/sites/<name>.zip using maximum compression.
  • HTTP delivery: The Express static middleware serves completed ZIPs directly, with the client receiving the final filename via the token-specific Socket.io channel.

Frequently Asked Questions

How does the server track individual download progress?

The server uses a token-based correlation system where each client request generates a unique identifier. The server emits progress events to a Socket.io channel named after this token (io.emit(data.token, {progress: txt})), ensuring that multiple concurrent downloads receive their specific status updates without interference.

What wget flags does the mirror process use and why?

The workflow executes wget with mkEpnp flags: -m enables mirror mode for recursive downloading, -k converts links for offline viewing, -E preserves HTML file extensions, -p downloads all prerequisites, and -np prevents ascending to parent directories. The --no-if-modified-since flag ensures fresh downloads regardless of server caching headers.

Where are the generated ZIP files stored and how are they served?

Completed archives are written to ./public/sites/<filename>.zip. The Express application serves this directory statically (configured in app.js via express.static), allowing clients to download files directly via standard HTTP requests to /sites/<filename>.zip once they receive the completion notification via WebSocket.

What happens if the client disconnects during the download?

The Socket.io disconnect handler in socket/socket.js attempts to kill the running wget child process and abort the archiver operation. Additionally, wget/index.js implements removePartiallyDownloadedFiles to clean up incomplete directory structures, preventing storage bloat from abandoned requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →