# How the Server Extracts and Determines Website Folder Names from URLs

> Discover how servers extract website folder names from URLs using a two-stage fallback. Learn hostname extraction from stderr and URL parsing for accurate folder determination.

- Repository: [Ahmed Ibrahim/Website-downloader](https://github.com/AhmadIbrahiim/Website-downloader)
- Tags: how-to-guide
- Published: 2026-07-08

---

**The server uses a two-stage fallback strategy: first capturing the hostname from wget's stderr output during the download process, then falling back to the `getWebsiteFolderName` helper function that normalizes and parses the URL to extract the hostname and port.**

When mirroring websites with `wget`, accurately mapping the source URL to the local filesystem folder is essential for organizing downloads. In the `AhmadIbrahiim/Website-downloader` repository, the Node.js server implements a robust extraction mechanism that handles protocol-less URLs, non-standard ports, and real-time process monitoring to determine the exact folder name created on disk.

## The Two-Stage Detection Strategy

### Stage 1: Real-Time Hostname Extraction from wget stderr

According to the implementation in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js), the server spawns a `wget` process with the flags `-mkEpnp` and immediately begins monitoring the **stderr** stream. The code watches for output lines matching the pattern:

```

Resolving <hostname> (

```

When this line appears, the server extracts `<hostname>` and stores it in the local variable `website`. This live capture ensures the folder name matches exactly what `wget` resolves during the download process.

### Stage 2: Fallback URL Parsing with getWebsiteFolderName

If the stderr monitoring fails to capture a hostname (for example, if the resolving line never appears or the output format differs), the server falls back to parsing the original request URL. The logic in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js) implements a simple OR condition:

```javascript
const websiteFolder = website || getWebsiteFolderName(data.website);

```

This ensures that even if real-time extraction fails, the server can still derive the correct folder name from the user-provided URL.

## Deep Dive into getWebsiteFolderName Implementation

The `getWebsiteFolderName` function in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js) handles URL normalization and extraction through three distinct operations:

### Protocol Normalization

Before parsing, the function ensures the URL has a valid scheme to satisfy the Node.js `URL` constructor:

```javascript
const normalizedUrl = /^https?:\/\//i.test(websiteUrl)
    ? websiteUrl
    : `http://${websiteUrl}`;

```

If the input lacks `http://` or `https://`, the code automatically prepends `http://` to prevent parsing errors.

### Hostname and Port Extraction

Using the built-in `URL` class, the function extracts the hostname and conditionally includes the port:

```javascript
const parsedUrl = new URL(normalizedUrl);

return parsedUrl.port
    ? `${parsedUrl.hostname}:${parsedUrl.port}`
    : parsedUrl.hostname;

```

When a non-standard port is present (e.g., `:8080`), the function returns `hostname:port` format. For standard ports (80/443), only the hostname is returned.

### Error Handling

Any parsing exceptions result in an empty string return:

```javascript
catch (error) {
    return "";
}

```

This empty string propagates through the system and triggers an error message sent back to the client before the archiving step begins.

## Code Examples and Edge Cases

The extraction logic handles various URL formats consistently:

```javascript
// Example 1: URL with explicit protocol
const folder1 = getWebsiteFolderName('https://my-site.org');
// Returns: 'my-site.org'

// Example 2: URL without protocol (gets normalized)
const folder2 = getWebsiteFolderName('my-site.org');
// Returns: 'my-site.org'

// Example 3: URL with non-standard port
const folder3 = getWebsiteFolderName('http://my-site.org:3000');
// Returns: 'my-site.org:3000'

```

## Integration with the Archiving Pipeline

Once determined, the `websiteFolder` string is passed directly to the archiving module. As shown in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js), the server calls:

```javascript
archiver(websiteFolder, io, data)

```

This function, implemented in [`archiver/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/archiver/index.js), creates the downloadable zip archive containing the mirrored website content. The folder name must match exactly the directory structure created by `wget` on the filesystem to ensure the archive packages the correct files.

## Summary

- The server prioritizes real-time extraction from wget's stderr output, capturing the hostname from the "Resolving" line during the download process
- The `getWebsiteFolderName` function in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js) provides a reliable fallback using Node.js's built-in `URL` parser with automatic protocol normalization
- URLs without schemes are automatically prefixed with `http://` before parsing to ensure compatibility with the `URL` constructor
- Non-standard ports are preserved in the folder name using the `hostname:port` format, while standard ports yield only the hostname
- Parsing errors return empty strings, which trigger client-side error messages before the archiving step in [`archiver/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/archiver/index.js)

## Frequently Asked Questions

### What happens if wget doesn't output a "Resolving" line?

If the stderr stream doesn't contain the expected "Resolving" pattern, the `website` variable remains undefined, and the server falls back to the `getWebsiteFolderName` function to parse the original URL from the request data.

### How does the server handle URLs without http:// or https:// prefixes?

The `getWebsiteFolderName` function automatically prepends `http://` to any URL lacking a protocol scheme using the regex `/^https?:\/\//i`, ensuring the Node.js `URL` constructor can parse it successfully without throwing errors.

### Does the folder name include the port number?

Yes, when a non-standard port is specified in the URL (e.g., `example.com:8080`), the function returns `hostname:port` format. Standard ports (80 for HTTP, 443 for HTTPS) are omitted by the `URL` class parser, so only the hostname is returned in those cases.

### Where is the extracted folder name used after determination?

According to the source code in [`wget/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/wget/index.js), the extracted `websiteFolder` string is passed directly to the `archiver` function imported from [`archiver/index.js`](https://github.com/AhmadIbrahiim/Website-downloader/blob/main/archiver/index.js), which creates the downloadable zip file containing the mirrored website content.