# Download Workflow from Request to ZIP Delivery: Technical Architecture of the Website Downloader

> Explore the Website Downloader architecture. Learn how Socket.io, wget, and an archiver module create a ZIP delivery pipeline from client request to file download.

- Repository: [Ahmed Ibrahim/Website-downloader](https://github.com/AhmadIbrahiim/Website-downloader)
- Tags: architecture
- Published: 2026-07-08

---

**The download workflow in AhmadIbrahiim/Website-downloader implements a Socket.io-driven pipeline where client requests trigger a wget child process to mirror websites, with stderr progress streamed back to the client in real-time before the archiver module compresses the result into a ZIP file stored in `public/sites/` for HTTP retrieval.**

The repository provides a Node.js service that converts live websites into offline ZIP archives. Understanding the download workflow from request to ZIP delivery reveals how WebSocket events, child processes, and file system operations coordinate to deliver complete website mirrors with real-time progress updates.

## Client Initiates Download via WebSocket

The workflow begins when the browser opens a Socket.io connection and emits a **`request`** event containing a unique **`token`** (to correlate responses) and the target **`website`** URL.

```javascript
// Client-side implementation
const socket = io();
const token = Date.now();

socket.emit('request', {
  token,
  website: 'https://example.com'
});

socket.on(token, data => {
  console.log('Progress:', data.progress);
  if (data.file) {
    window.location = `/sites/${data.file}.zip`;
  }
});

```

The client listens on a channel named after the token to receive real-time progress updates and the final download URL.

## Server Receives and Routes the Request

**File:** [`socket/socket.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/socket/socket.js)

The server-side Socket.io listener captures the request and immediately forwards it to the download worker:

```javascript
socket.on('request', function (data) {
  console.log("Request connection received %s", data.token);
  wget(io, data);  // Delegates to wget/index.js
});

```

The global `io` instance (Socket.io server) and request data are passed to the `wget()` function, establishing the connection between the real-time messaging layer and the file system operations.

## wget Mirrors the Website with Real-Time Progress Streaming

**File:** [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js)

The download workflow spawns a child process using Node.js `exec()` with specific wget flags designed for complete website mirroring:

```javascript
const child = exec(`wget -mkEpnp --no-if-modified-since ${data.website}`);

```

The **`-mkEpnp`** flags configure wget to create a recursive mirror, convert links for offline use, preserve file extensions, and avoid ascending to parent directories during the crawl.

### Progress Transmission

All stderr output streams back to the client via Socket.io events, enabling live progress bars in the UI:

```javascript
child.stderr.on("data", (response) => {
  const txt = response.toString();
  // Extract domain name from "Resolving..." output
  if (!website) {
    const m = txt.match(/Resolving\s+([^\s]+)\s+\(/);
    if (m) website = m[1];
  }
  io.emit(data.token, { progress: txt });
});

```

### Folder Detection and Handoff

When wget completes, the `close` event handler determines the downloaded folder name and triggers the archiving phase:

```javascript
child.stderr.on('close', () => {
  const websiteFolder = website || getWebsiteFolderName(data.website);
  io.emit(data.token, {progress: "Converting"});
  archiver(websiteFolder, io, data);  // Transition to zipping stage
});

```

The `getWebsiteFolderName()` utility extracts the hostname (or host:port) from the original URL if wget's output parsing failed to identify the domain.

## Archiving the Downloaded Directory into ZIP

**File:** [`archiver/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/archiver/index.js)

The archiver module creates a compressed ZIP file with maximum compression settings:

```javascript
var output = fs.createWriteStream("./public/sites/" + file + '.zip');
var archive = archiver('zip', { zlib: { level: 9 } });

output.on('close', () => {
  console.log(archive.pointer() + ' total bytes');
  io.emit(data.token, { progress: "Completed", file });
});

archive.pipe(output);
archive.directory('./' + file, false);  // Include entire site folder
archive.finalize();

```

The `level: 9` zlib setting applies maximum compression. The archive pipes to `./public/sites/<file>.zip`, making it accessible via HTTP because [`app.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/app.js) serves the `public` directory statically using `express.static()`.

## Error Handling and Cleanup

The download workflow includes mechanisms for handling client disconnections and partial failures. In [`socket/socket.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/socket/socket.js), disconnect events attempt to terminate running wget processes and abort archiver operations. The [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js) module implements `removePartiallyDownloadedFiles` to clean up incomplete downloads when processes terminate unexpectedly.

Both wget and archiver forward non-critical warnings while throwing on fatal errors, preventing silent failures in the pipeline.

## Summary

- **WebSocket initiation**: Clients emit `request` events with unique tokens and target URLs via Socket.io.
- **Server routing**: [`socket/socket.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/socket/socket.js) receives requests and delegates to [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js).
- **Mirroring process**: `wget` spawns with `-mkEpnp` flags to create offline mirrors while streaming stderr progress back to the client.
- **Folder detection**: The system parses wget output or extracts hostnames to identify the downloaded directory.
- **ZIP creation**: [`archiver/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/archiver/index.js) compresses the folder into `public/sites/<name>.zip` using maximum compression.
- **HTTP delivery**: The Express static middleware serves completed ZIPs directly, with the client receiving the final filename via the token-specific Socket.io channel.

## Frequently Asked Questions

### How does the server track individual download progress?

The server uses a **token-based correlation system** where each client request generates a unique identifier. The server emits progress events to a Socket.io channel named after this token (`io.emit(data.token, {progress: txt})`), ensuring that multiple concurrent downloads receive their specific status updates without interference.

### What wget flags does the mirror process use and why?

The workflow executes wget with **`mkEpnp`** flags: `-m` enables mirror mode for recursive downloading, `-k` converts links for offline viewing, `-E` preserves HTML file extensions, `-p` downloads all prerequisites, and `-np` prevents ascending to parent directories. The **`--no-if-modified-since`** flag ensures fresh downloads regardless of server caching headers.

### Where are the generated ZIP files stored and how are they served?

Completed archives are written to **`./public/sites/<filename>.zip`**. The Express application serves this directory statically (configured in [`app.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/app.js) via `express.static`), allowing clients to download files directly via standard HTTP requests to `/sites/<filename>.zip` once they receive the completion notification via WebSocket.

### What happens if the client disconnects during the download?

The Socket.io disconnect handler in [`socket/socket.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/socket/socket.js) attempts to kill the running wget child process and abort the archiver operation. Additionally, [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js) implements `removePartiallyDownloadedFiles` to clean up incomplete directory structures, preventing storage bloat from abandoned requests.