How the Website-Downloader Application Manages Temporary Files and Disk Space Cleanup

The Website-downloader application stores mirrored sites in temporary folders named after the target host, archives them into ZIP files in public/sites/, and only deletes those temporary folders when the download process receives a SIGTERM signal, leaving successful downloads on disk until manual or scheduled removal.

The open-source repository AhmadIbrahiim/Website-downloader provides a Node.js application that downloads entire websites using wget and packages them for user retrieval. Understanding how this tool handles temporary files and disk space cleanup is critical for maintaining server storage, as the default behavior leaves completed downloads on the filesystem while only cleaning up aborted attempts.

Temporary File Creation During Site Mirroring

When a user initiates a download via the socket.io interface, the application spawns a wget process that mirrors the target site into a local directory.

Host-Based Folder Naming

In wget/index.js, the temporary folder name is derived from the requested URL. Lines 37–38 assign the folder name based on the website variable (or via the getWebsiteFolderName function), creating a directory named after the target host (e.g., example.com):

// wget/index.js (lines 37-38)
var website = getWebsiteFolderName(website);
// wget process writes to ./website/

This folder resides in the application root and serves as the working directory for the wget process, accumulating all downloaded assets, HTML files, and subdirectories until the mirroring completes or fails.

Archiving and Permanent Storage

Once the wget process finishes successfully, the application transfers the content from the temporary folder to a permanent archive location.

Creating ZIP Archives from Temporary Directories

The archiver/index.js module handles the compression. Lines 7–9 configure the output stream to write the final ZIP file to public/sites/:

// archiver/index.js (lines 7-9)
const output = fs.createWriteStream(__dirname + '/public/sites/' + file + '.zip');
const archive = archiver('zip', {
  zlib: { level: 9 }
});

At line 48, the archiver adds the entire temporary directory to the ZIP using archive.directory('./'+file,false);, which recursively includes all downloaded content. After the archive finalizes, the ZIP file remains in public/sites/ for user download, but the source temporary folder in the root directory is not automatically deleted.

Cleanup Logic for Aborted Downloads

The application implements selective cleanup that only triggers during abnormal termination scenarios.

SIGTERM Handling and Partial Download Removal

If the user aborts the request or the process is killed, the child process emits an 'exit' event. The handler at lines 48–54 in wget/index.js checks for the SIGTERM signal and invokes removePartiallyDownloadedFiles(website):

// wget/index.js (lines 48-54)
child.on('exit', (code, signal) => {
  if (signal === 'SIGTERM') {
    console.log('Process terminated');
    removePartiallyDownloadedFiles(website);
  }
});

The removePartiallyDownloadedFiles Function

The cleanup function, defined at lines 56–64, uses fs.rmdir with the recursive: true option to delete the entire temporary folder and its contents:

// wget/index.js (lines 56-64)
function removePartiallyDownloadedFiles(website) {
  fs.rmdir('./' + website, { recursive: true }, (err) => {
    if (err) {
      console.error(err);
      return;
    }
    console.log('Removed partially downloaded files');
  });
}

This prevents orphaned data from accumulating when downloads fail or are cancelled, but it does not execute after a successful download-and-archive cycle.

Disk Space Management Gaps and Recommendations

The current implementation creates a storage liability: successful downloads leave the temporary folder intact alongside the permanent ZIP archive. Running many downloads causes the root directory to accumulate host-named folders (e.g., example.com, another-site.org) indefinitely.

Implementing a Periodic Cleanup Routine

To prevent uncontrolled disk usage, add a scheduled cleanup job that removes temporary folders older than a specified threshold. Place this logic in your server entry file (e.g., app.js):

const fs = require('fs');
const path = require('path');

// Remove temporary download folders older than 24 hours
setInterval(() => {
  const root = path.join(__dirname, '..');
  fs.readdir(root, (err, files) => {
    if (err) return;
    files.forEach(f => {
      const dir = path.join(root, f);
      fs.stat(dir, (e, stats) => {
        if (!e && stats.isDirectory() && Date.now() - stats.mtimeMs > 24 * 60 * 60 * 1000) {
          fs.rmdir(dir, { recursive: true }, err => {
            if (!err) console.log(`Cleaned up old folder: ${f}`);
          });
        }
      });
    });
  });
}, 6 * 60 * 60 * 1000); // Run twice daily

Alternatively, configure a system cron job to remove directories older than 24 hours from the application root, or modify archiver/index.js to delete the source folder immediately after confirming the ZIP archive is valid.

Summary

  • Temporary folders are created in the application root using the target hostname (e.g., example.com) as the directory name in wget/index.js lines 37–38.
  • Archiving occurs via archiver/index.js lines 7–9 and 48, which write ZIP files to public/sites/ but leave the source folder intact.
  • Cleanup only runs on abort: The removePartiallyDownloadedFiles function at wget/index.js lines 56–64 executes only when the wget process exits with a SIGTERM signal.
  • Successful downloads persist: Temporary folders remain on disk after archiving completes, requiring manual deletion or a periodic cleanup routine to prevent disk space exhaustion.

Frequently Asked Questions

When does the application delete temporary files?

The application deletes temporary files only when the wget process is terminated by a SIGTERM signal. According to the exit handler in wget/index.js lines 48–54, this triggers removePartiallyDownloadedFiles(), which recursively deletes the host-named folder. Successful downloads do not trigger automatic cleanup.

Where are the downloaded website archives stored?

Completed archives are stored in the public/sites/ directory relative to the archiver/index.js file. The archiver module uses fs.createWriteStream to write ZIP files to this location at lines 7–9, making them available for user download while the temporary source data remains in the application root.

Why does disk space fill up even after downloads complete?

Disk space fills up because the temporary folder created by wget (e.g., ./example.com) is not deleted after the archiver successfully creates the ZIP file. The removePartiallyDownloadedFiles function only runs during SIGTERM aborts, not after normal completion, leaving the original mirrored site data on disk alongside the compressed archive.

How can I prevent orphaned temporary folders from accumulating?

Implement a periodic cleanup mechanism that deletes folders older than 24 hours from the application root directory, or modify the archiver to delete the source folder after verifying the ZIP archive integrity. The repository does not include automatic cleanup for successful downloads, so external maintenance via cron jobs or the setInterval pattern shown above is necessary for production deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →