How the Server Extracts and Determines Website Folder Names from URLs
The server uses a two-stage fallback strategy: first capturing the hostname from wget's stderr output during the download process, then falling back to the getWebsiteFolderName helper function that normalizes and parses the URL to extract the hostname and port.
When mirroring websites with wget, accurately mapping the source URL to the local filesystem folder is essential for organizing downloads. In the AhmadIbrahiim/Website-downloader repository, the Node.js server implements a robust extraction mechanism that handles protocol-less URLs, non-standard ports, and real-time process monitoring to determine the exact folder name created on disk.
The Two-Stage Detection Strategy
Stage 1: Real-Time Hostname Extraction from wget stderr
According to the implementation in wget/index.js, the server spawns a wget process with the flags -mkEpnp and immediately begins monitoring the stderr stream. The code watches for output lines matching the pattern:
Resolving <hostname> (
When this line appears, the server extracts <hostname> and stores it in the local variable website. This live capture ensures the folder name matches exactly what wget resolves during the download process.
Stage 2: Fallback URL Parsing with getWebsiteFolderName
If the stderr monitoring fails to capture a hostname (for example, if the resolving line never appears or the output format differs), the server falls back to parsing the original request URL. The logic in wget/index.js implements a simple OR condition:
const websiteFolder = website || getWebsiteFolderName(data.website);
This ensures that even if real-time extraction fails, the server can still derive the correct folder name from the user-provided URL.
Deep Dive into getWebsiteFolderName Implementation
The getWebsiteFolderName function in wget/index.js handles URL normalization and extraction through three distinct operations:
Protocol Normalization
Before parsing, the function ensures the URL has a valid scheme to satisfy the Node.js URL constructor:
const normalizedUrl = /^https?:\/\//i.test(websiteUrl)
? websiteUrl
: `http://${websiteUrl}`;
If the input lacks http:// or https://, the code automatically prepends http:// to prevent parsing errors.
Hostname and Port Extraction
Using the built-in URL class, the function extracts the hostname and conditionally includes the port:
const parsedUrl = new URL(normalizedUrl);
return parsedUrl.port
? `${parsedUrl.hostname}:${parsedUrl.port}`
: parsedUrl.hostname;
When a non-standard port is present (e.g., :8080), the function returns hostname:port format. For standard ports (80/443), only the hostname is returned.
Error Handling
Any parsing exceptions result in an empty string return:
catch (error) {
return "";
}
This empty string propagates through the system and triggers an error message sent back to the client before the archiving step begins.
Code Examples and Edge Cases
The extraction logic handles various URL formats consistently:
// Example 1: URL with explicit protocol
const folder1 = getWebsiteFolderName('https://my-site.org');
// Returns: 'my-site.org'
// Example 2: URL without protocol (gets normalized)
const folder2 = getWebsiteFolderName('my-site.org');
// Returns: 'my-site.org'
// Example 3: URL with non-standard port
const folder3 = getWebsiteFolderName('http://my-site.org:3000');
// Returns: 'my-site.org:3000'
Integration with the Archiving Pipeline
Once determined, the websiteFolder string is passed directly to the archiving module. As shown in wget/index.js, the server calls:
archiver(websiteFolder, io, data)
This function, implemented in archiver/index.js, creates the downloadable zip archive containing the mirrored website content. The folder name must match exactly the directory structure created by wget on the filesystem to ensure the archive packages the correct files.
Summary
- The server prioritizes real-time extraction from wget's stderr output, capturing the hostname from the "Resolving" line during the download process
- The
getWebsiteFolderNamefunction inwget/index.jsprovides a reliable fallback using Node.js's built-inURLparser with automatic protocol normalization - URLs without schemes are automatically prefixed with
http://before parsing to ensure compatibility with theURLconstructor - Non-standard ports are preserved in the folder name using the
hostname:portformat, while standard ports yield only the hostname - Parsing errors return empty strings, which trigger client-side error messages before the archiving step in
archiver/index.js
Frequently Asked Questions
What happens if wget doesn't output a "Resolving" line?
If the stderr stream doesn't contain the expected "Resolving" pattern, the website variable remains undefined, and the server falls back to the getWebsiteFolderName function to parse the original URL from the request data.
How does the server handle URLs without http:// or https:// prefixes?
The getWebsiteFolderName function automatically prepends http:// to any URL lacking a protocol scheme using the regex /^https?:\/\//i, ensuring the Node.js URL constructor can parse it successfully without throwing errors.
Does the folder name include the port number?
Yes, when a non-standard port is specified in the URL (e.g., example.com:8080), the function returns hostname:port format. Standard ports (80 for HTTP, 443 for HTTPS) are omitted by the URL class parser, so only the hostname is returned in those cases.
Where is the extracted folder name used after determination?
According to the source code in wget/index.js, the extracted websiteFolder string is passed directly to the archiver function imported from archiver/index.js, which creates the downloadable zip file containing the mirrored website content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →