How AhmadIbrahiim/Website-downloader Handles Concurrent Download Requests from Multiple Users

The Website-downloader application handles concurrent download requests by spawning isolated Node.js child processes for each Socket.io connection, ensuring every user receives a dedicated OS process with token-based progress tracking.

The AhmadIbrahiim/Website-downloader repository provides a Node.js-based solution for downloading complete websites. When multiple users initiate downloads simultaneously, the system leverages Socket.io connections and Node.js child processes to maintain strict isolation between sessions. This architecture ensures that concurrent download requests operate independently without shared state or resource conflicts.

Architecture Overview

The system implements a per-connection process isolation pattern. When a user opens the web interface, the server establishes a Socket.io connection that manages the entire lifecycle of that download session. Unlike traditional request-response models that might queue operations, this implementation immediately forks new operating system processes for each incoming request.

According to the source code in socket/socket.js, the server listens for incoming request events and invokes the wget function defined in wget/index.js. This function spawns a new child process using child_process.exec, creating a separate shell environment for the wget command.

How Concurrent Requests Are Processed

Per-Socket Process Isolation

Each download request triggers the creation of two distinct child processes:

  1. Primary download process: Executes the wget command to mirror the target website
  2. Archival process: Compresses the downloaded content into a ZIP file

In wget/index.js (lines 7-22), the exec function spawns the wget process with the command:

const child = exec(`wget -mkEpnp --no-if-modified-since ${data.website}`);

Because this executes within the context of a specific Socket.io connection, every concurrent user receives their own dedicated OS process. The system does not implement a shared worker pool or queue; instead, it relies on the operating system's process scheduler to handle parallel execution.

Token-Based Message Routing

To prevent cross-talk between concurrent downloads, the system uses a unique token generated by the client. As shown in socket/socket.js (lines 6-8), each request event carries a token and website URL:

socket.emit('request', { token, website: 'https://example.com' });

The server uses this token to route progress updates exclusively to the requesting client. In wget/index.js (lines 33-34), stderr data from the wget process emits to the specific token channel:

child.stderr.on('data', response => {
  io.emit(data.token, { progress: response.toString() });
});

This ensures that when multiple users download simultaneously, each client receives only their own progress updates.

Archival Concurrency

After the wget process completes, the system spawns an additional archiver process. The archiver function in wget/index.js (lines 44-46) initializes this compression step, which runs independently and reports completion through the same token-based channel.

Automatic Cleanup on Disconnection

The system implements graceful cleanup to prevent resource exhaustion when users disconnect. In socket/socket.js (lines 11-22), the server listens for disconnect events and terminates associated child processes:

socket.on('disconnect', () => {
  if (socket.wgetProcess) socket.wgetProcess.kill();
  if (socket.archiverProcess) socket.archiverProcess.abort();
});

This prevents "orphaned" download processes from consuming server resources when a client closes their browser or loses network connectivity during active concurrent sessions.

Implementation Examples

Client-Side Request Initiation

// Browser environment with Socket.io client
const token = Date.now().toString();  // Unique identifier for this session
socket.emit('request', { 
  token, 
  website: 'https://example.com' 
});

// Listen for progress updates specific to this token
socket.on(token, data => {
  console.log('Download progress:', data.progress);
  if (data.file) {
    // Trigger download of the final ZIP archive
    window.location = `/public/sites/${data.file}.zip`;
  }
});

Server-Side Process Spawning

// wget/index.js
const { exec } = require('child_process');

function wget(io, data) {
  // Spawn isolated wget process
  const wgetProcess = exec(`wget -mkEpnp --no-if-modified-since ${data.website}`);
  
  // Stream progress back to client using token as event name
  wgetProcess.stderr.on('data', response => {
    io.emit(data.token, { progress: response.toString() });
  });
  
  // Trigger archival after download completes
  wgetProcess.stderr.on('close', () => {
    const folder = getWebsiteFolderName(data.website);
    archiver(folder, io, data);  // Starts zip process
  });
}

Summary

  • Process-per-connection architecture: Each concurrent download request spawns independent wget and archiver child processes via child_process.exec
  • Token-based isolation: Unique client-generated tokens ensure progress updates route to the correct user during concurrent sessions
  • Automatic resource cleanup: The disconnect event handler terminates child processes to prevent server resource exhaustion
  • No shared state: Downloads operate in isolated OS processes, eliminating race conditions between multiple users

Frequently Asked Questions

What happens if two users download the same website simultaneously?

Each user receives their own dedicated child process. Because the system stores process references per Socket.io connection (see socket/socket.js), two users downloading the same URL operate in completely isolated environments with separate file system writes and network streams.

Is there a limit to how many concurrent downloads the system can handle?

The practical limit depends on the host machine's available memory and CPU cores, as each download spawns separate OS processes. The system does not implement an application-level queue or rate limiter; concurrency is constrained only by operating system resources and ulimit settings for maximum open files and processes.

How does the system prevent progress updates from mixing between users?

The client generates a unique token (typically Date.now().toString()) when initiating the request. The server uses this token as the Socket.io event name when emitting progress updates from wget/index.js. This ensures that even with hundreds of concurrent connections, each browser receives only the updates associated with its specific token.

What happens to active downloads if a user closes their browser?

The disconnect event handler in socket/socket.js attempts to kill the associated wgetProcess and abort the archiverProcess. This prevents orphaned processes from continuing to consume bandwidth and CPU cycles after the client connection drops, protecting server resources during high-concurrency scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →