Wigolo's Tiered Fetch Router Architecture: How It Handles Anti-Bot Challenges
Wigolo's SmartRouter automatically escalates through TLS impersonation, HTTP, and Playwright browser tiers based on domain heuristics and real-time challenge detection to bypass anti-bot protections.
Wigolo (KnockOutEZ/wigolo) implements a sophisticated tiered fetch router that dynamically selects the optimal fetching strategy for each request. The architecture prioritizes speed and efficiency while maintaining robust capabilities for penetrating modern anti-bot systems like Cloudflare and DataDome.
The Four-Tier Fetch Strategy
The SmartRouter defined in src/fetch/router.ts orchestrates four distinct fetching tiers:
| Tier | Implementation | Trigger Conditions |
|---|---|---|
| TLS Impersonation | tlsFetch from src/fetch/tls-tier.ts |
Global TLS setting enabled, domain marked in ANTI_BOT_TLS_DOMAINS, or user-provided WIGOLO_TLS_DOMAINS |
| HTTP Client | httpClient.fetch |
Default path for standard requests without anti-bot signals |
| Playwright Browser | fetchWithPlaywright from src/fetch/playwright-tier.ts |
JavaScript rendering required, actions specified, anti-bot signals detected, or SPA domains |
| Escape Hatch | External solver/hosted reader | Configured fallback after ChallengeBlockedError in browser tier |
Router Initialization and Configuration
The router accepts configuration through the RouterFetchOptions interface, allowing per-request customization of the tier selection logic:
// src/fetch/router.ts:84-98
export interface RouterFetchOptions {
renderJs?: 'auto' | 'always' | 'never';
useAuth?: boolean;
headers?: Record<string, string>;
screenshot?: boolean;
actions?: BrowserAction[];
force_refresh?: boolean;
mode?: Mode;
conditionalHeaders?: { ifNoneMatch?: string; ifModifiedSince?: string };
signal?: AbortSignal;
}
During construction (lines 55-70), the SmartRouter wires default implementations for each tier, binding this.tlsFetcher = tlsFetch and this.playwrightFetcher = fetchWithPlaywright to enable the escalation pipeline.
Domain-Level Heuristics and Routing Decisions
Before executing any network request, the router evaluates domain-specific rules defined in src/fetch/router.ts to pre-select the appropriate tier.
Known SPA Domains
Modern single-page applications often return empty HTML shells on static requests. The router maintains a curated list to force Playwright on first visit:
// src/fetch/router.ts:48-64
const KNOWN_SPA_DOMAINS = new Set<string>([
'react.dev', 'nextjs.org', 'vuejs.org', // additional SPA frameworks
]);
If the target domain exists in this set, the router bypasses HTTP and proceeds directly to the Playwright tier, avoiding the latency of a failed static fetch.
Anti-Bot TLS Domains
Certain domains block standard HTTP connections entirely, requiring TLS-level impersonation to establish any connection:
// src/fetch/router.ts:66-75
const ANTI_BOT_TLS_DOMAINS = new Set<string>([
'stackoverflow.com', 'serverfault.com', // sites that timeout before HTTP response
]);
These domains trigger TLS-first routing regardless of global configuration settings, as the TLS impersonation tier can break through connection-level anti-bot blocks that would otherwise terminate the request.
The Main Fetch Execution Flow
The core routing logic resides in the overloaded fetch method starting at line 887. The implementation first evaluates early-exit conditions before entering the default auto mode:
// src/fetch/router.ts:996-1003
if (mode === 'stealth') { /* static → Playwright escalation */ }
if (mode === 'cache') { /* HTTP-only, tiny timeout */ }
if (actions?.length) { /* actions force Playwright */ }
if (renderJs === 'always' || useAuth) { /* explicit Playwright */ }
if (renderJs === 'never') { /* HTTP-only, no escalation */ }
if (looksLikeBinaryDownload(url)) { /* force HTTP for PDFs, ZIPs */ }
When operating in auto mode, the router executes the following sequence:
- Rate-limit back-off check – If the host is currently in a back-off window (
backoffWindowResult), the router returns a stage error immediately (lines 79-82) - Domain learning check – If previous requests promoted the domain (
stats.preferPlaywright), the router jumps tobrowserOrHttpForBinary(lines 88-91) - TLS-first calculation –
tryTlsFirstis computed from configuration, learned preferences, and the anti-bot allowlist (lines 94-108) - Tier execution – The router attempts TLS if indicated, then falls back to HTTP, finally escalating to Playwright if anti-bot signals appear
Anti-Bot Detection and Escalation Strategies
The router implements multiple detection mechanisms to identify and respond to anti-bot challenges in src/fetch/router.ts.
Bare 403 with Rotated User-Agent
When encountering a 403 status without a challenge body, the router performs a single retry with a stealth User-Agent:
// src/fetch/router.ts:1042-1050
if (result.statusCode === 403 && !isChallengeResponse(...)) {
const rotatedUa = resolveStealthUA();
// ...
const retried = await this.httpClientFetch(...);
if (!isAntiBotSignal(retried.statusCode, retried.html)) {
// Success on retry
}
}
Challenge Shell Detection at 2xx
Anti-bot systems sometimes return 200 status codes with challenge JavaScript shells. The isChallengeShell detection triggers tier escalation:
// src/fetch/router.ts:1054-1070
if (is2xxShell) {
if (tlsUsable) { /* attempt TLS bypass */ }
stats.preferPlaywright = true;
return this.browserFetch(url, {
// ...,
stealth: stealthForBrowser(config, { antiBotEscalation: true }),
fallback: /* ... */
});
}
Upon detection, the domain is marked with stats.preferPlaywright = true to avoid future static attempts.
Anti-Bot Signal Handling (403/503/429)
The isAntiBotSignal function identifies challenge responses by status code (≥403 or 429) combined with challenge body detection. When triggered with usable TLS:
// Conceptual flow from lines 115-128
if (isAntiBotSignal(status, html) && tlsUsable) {
const tlsResult = await this.tryTlsTier(url, options);
if (!tlsResult.ok) {
return this.browserFetch(url, options); // Escalate to Playwright
}
}
Browser-Tier Challenge Blocking
When the Playwright tier encounters a hardened challenge it cannot solve, it throws ChallengeBlockedError. The router catches this and attempts configured escape-hatch services:
// src/fetch/router.ts:1110-1122
if (err instanceof ChallengeBlockedError) {
const cleared = await this.tryEscapeHatch(url, browserOptions.signal);
if (cleared) return cleared;
return { error: err.code, /* ... */ };
}
Challenge Shell Guarding
Both HTTP and browser results pass through guardChallengeShell (lines 71-80) to prevent leaking raw challenge HTML to callers:
// src/fetch/router.ts:71-80
private guardChallengeShell(raw) {
if (isChallengeResponse(raw.statusCode, raw.html, raw.headers)) {
const err = new ChallengeBlockedError(raw.url);
return { error: err.code, /* ... */ };
}
return raw;
}
Performance Optimizations and Caching
The tiered router includes several mechanisms to minimize redundant anti-bot solving:
- Clearance Reuse – Cookies from previously solved challenges are injected into TLS and HTTP requests via
withClearanceHeader(lines 48-62), eliminating the need to re-solve challenges for subsequent requests to the same host - Back-off Queue – When receiving 429 responses, the router invokes
recordBackoffto park the host for a bounded period, preventing hammering of rate-limited origins (lines 71-82) - TLS Persistence – Successful TLS attempts are recorded via
tlsPersistence.getPreferTls(domain)to promote future TLS-first routing for that domain (lines 94-108) - Binary Download Shortcuts – URLs ending in binary extensions (PDF, ZIP) bypass Playwright entirely to avoid download corruption (lines 1069-1072)
Implementation Example
The following example demonstrates initializing the router with HTTP and browser pool dependencies:
import { SmartRouter } from './fetch/router.js';
import { createHttpClient } from './http-client.js';
import { createBrowserPool } from './browser-pool.js';
const router = new SmartRouter({
httpClient: createHttpClient(),
browserPool: createBrowserPool(),
// Optional: custom TLS fetcher, clearance store, etc.
});
(async () => {
// Auto mode handles anti-bot challenges transparently
const result = await router.fetch('https://stackoverflow.com/questions/12345');
console.log(result.html?.slice(0, 200));
})();
When targeting Stack Overflow, the router automatically selects TLS-first routing due to its presence in ANTI_BOT_TLS_DOMAINS, then falls back through HTTP to Playwright only if Cloudflare presents an interactive challenge.
Summary
- Wigolo's
SmartRouterinsrc/fetch/router.tsimplements a four-tier architecture: TLS impersonation, HTTP, Playwright browser, and optional external solvers - Domain heuristics in
KNOWN_SPA_DOMAINSandANTI_BOT_TLS_DOMAINSenable intelligent pre-selection of the optimal tier - Real-time detection via
isAntiBotSignalandisChallengeShelltriggers automatic escalation from static to browser-based fetching - Challenge solutions are cached and reused across tiers via
src/fetch/clearance-reuse.tsto minimize redundant solving - The escape-hatch mechanism provides integration points for third-party solvers when browser automation fails
Frequently Asked Questions
What is Wigolo's tiered fetch router?
Wigolo's tiered fetch router is a request orchestration system implemented in src/fetch/router.ts that dynamically selects between TLS impersonation, HTTP, and Playwright browser tiers based on domain characteristics and real-time anti-bot detection. It automatically escalates through tiers when encountering blocks, ensuring maximum success rates while minimizing resource usage for simple requests.
How does the TLS impersonation tier bypass anti-bot protections?
The TLS tier (src/fetch/tls-tier.ts) uses tlsFetch to generate TLS fingerprints that mimic real browsers at the connection level. This bypasses network-layer blocks that identify automated traffic through TLS handshake characteristics, particularly effective for domains listed in ANTI_BOT_TLS_DOMAINS that timeout or reset connections from standard HTTP clients.
When does the router escalate to the Playwright browser tier?
The router escalates to Playwright (fetchWithPlaywright) when: (1) the domain is in KNOWN_SPA_DOMAINS, (2) renderJs is set to 'always' or 'auto' with detected JavaScript requirements, (3) browser actions are specified, (4) anti-bot signals are detected via isAntiBotSignal, or (5) a previous request marked the domain with preferPlaywright due to challenge detection.
How does the router optimize performance for repeated requests?
The router maintains domain-level state in src/cache/store.ts including clearance cookies (reused via withClearanceHeader), TLS preference flags (tlsPersistence), and back-off timers to avoid rate-limited hosts. Successful challenge solutions are cached and injected into subsequent HTTP/TLS requests, eliminating the latency of re-solving challenges for previously accessed domains.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →