# How Wigolo Handles Anti-Bot Challenges and SPA Pages in Its Fetch Router

> Discover how Wigolo's fetch router tackles anti-bot challenges and SPA pages using a multi-tier strategy from HTTP to Playwright browser automation.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-29

---

**Wigolo's fetch router implements a multi-tier escalation strategy that attempts plain HTTP first, falls back to TLS impersonation for known anti-bot domains, and finally escalates to Playwright browser automation when challenge shells or empty SPA content is detected.**

The `SmartRouter` class in [`src/fetch/router.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/router.ts) serves as the central dispatcher for all outbound requests in the [Wigolo](https://github.com/KnockOutEZ/wigolo) scraping framework. It intelligently routes traffic between three distinct fetching tiers—plain HTTP, TLS-impersonation, and Playwright-driven browsers—based on real-time detection of anti-bot challenges and Single-Page Application (SPA) rendering requirements.

## Detecting Anti-Bot Challenges

The router relies on signal detection helpers imported from [`src/fetch/tls-tier.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/tls-tier.ts) to identify when a target has deployed protective measures that require escalation beyond standard HTTP fetching.

### Signal Detection via Status Codes and Body Markers

Wigolo identifies anti-bot protections through two primary helper functions:

- **`isAntiBotSignal`** (lines 1020-1024 in [`src/fetch/tls-tier.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/tls-tier.ts)): Returns `true` for HTTP status codes **403**, **429**, or **503**, or when response bodies contain known challenge markers indicating rate limiting or blocking.

- **`isChallengeShell`** (lines 1059-1065): Classifies responses as *challenge interstitials*—such as Cloudflare's "Just a moment" or DataDome's "enable JavaScript" pages—even when the HTTP status code appears successful (2xx).

These functions allow the router to distinguish between genuine content and protective barriers regardless of the HTTP status returned.

### The Three-Tier Escalation Path

When the router detects a challenge signal, it follows a strict escalation sequence designed to minimize expensive browser launches:

1. **TLS-First Attempt**: If TLS impersonation is enabled (via `getConfig().tlsTier === 'on'` or when the domain exists in the anti-bot allow-list checked by `isAntiBotTlsDomain` at lines 106-108), the router invokes `tryTlsTier` to masquerade as a legitimate browser fingerprint at the TLS layer.

2. **Playwright Fallback**: If TLS fails, times out, or is unavailable, the router calls `browserFetch`, which wraps the request with `guardChallengeShell` (lines 886-896) to convert any remaining challenge HTML into a structured `blocked_by_challenge` error rather than returning raw interstitial markup.

3. **Error Structuring**: The `guardChallengeShell` wrapper ensures that even sophisticated challenges that slip past TLS detection are caught and returned as machine-readable errors instead of opaque HTML strings.

### Stealth Mode Optimization

When callers specify `mode: 'stealth'`, the router defers Playwright instantiation through a two-phase approach:

- First, it performs a lightweight static fetch using the HTTP client.
- It only escalates to Playwright if `shouldEscalate` determines the content is thin **or** if `isChallengeShell` detects a challenge interstitial.

This optimization prevents unnecessary browser launches for simple static sites while maintaining the capability to penetrate sophisticated anti-bot walls when required.

## Handling Single-Page Application (SPA) Pages

Modern SPAs render meaningful content only after JavaScript execution, requiring special routing logic beyond standard static HTML fetching.

### Pre-Configured SPA Domain Lists

The router maintains a hard-coded constant array `KNOWN_SPA_DOMAINS` (lines 52-64 in [`src/fetch/router.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/router.ts)) containing hosts that are heavily client-rendered. When a URL matches this list, the router bypasses HTTP entirely and routes directly to Playwright, eliminating the latency of failed static fetches on known JavaScript-heavy platforms like React documentation sites or Angular applications.

### Per-Domain Preference Learning

Wigolo implements intelligent caching through the `DomainStats` interface (lines 44-46), which tracks a `preferPlaywright` boolean flag for each unique host:

- **Initial Routing**: If a host exists in `KNOWN_SPA_DOMAINS`, the router sets `preferPlaywright = true` and immediately routes to the browser tier (see the conditional block at lines 1004-1006).

- **Dynamic Downgrade**: If a subsequent HTTP fetch for that domain returns non-empty HTML that passes the `contentAppearsEmpty` check, the router updates the statistics to set `preferPlaywright = false` (lines 1065-1068). This optimization prevents unnecessary browser launches for domains that have since implemented server-side rendering or static generation.

### Automatic Fallback for Unknown Domains

When `renderJs` is set to `'auto'` (the default configuration), the router employs a heuristic approach:

1. Attempt HTTP fetching first.
2. If the response triggers `contentAppearsEmpty` (indicating JavaScript-required content) or encounters a challenge shell, automatically escalate to Playwright (logic implemented at lines 1019-1035).

This automatic fallback balances fetch performance against rendering accuracy for unknown sites without requiring manual domain configuration.

## Practical Implementation Examples

The following examples demonstrate how to interact with Wigolo's fetch router for common use cases:

```typescript
// 1. Simple fetch – let the router decide (auto mode)
import { SmartRouter } from 'wigolo/src/fetch/router.js';

const router = new SmartRouter(/* httpClient & browserPool injected */);
const result = await router.fetch('https://react.dev/tutorial/tutorial.html');
// → Playwright is used because react.dev is in KNOWN_SPA_DOMAINS

```

```typescript
// 2. Force TLS-first for a domain that often times out (e.g., StackOverflow)
import { getConfig } from 'wigolo/src/config.js';

getConfig().tlsTier = 'on';          // enable TLS tier globally
getConfig().tlsDomains = ['stackoverflow.com']; // or rely on built-in list

const tlsFirst = await router.fetch('https://stackoverflow.com/q/123456');
/* → TLS tier is tried first; if it succeeds the content is returned,
   otherwise the router falls back to HTTP and then Playwright if needed. */

```

```typescript
// 3. Stealth mode – static fetch + optional Playwright escalation
const stealthResult = await router.fetch(
  'https://example.com/protected',
  { mode: 'stealth' }
);
// → Returns static HTML unless it is thin or a challenge shell,
//    in which case Playwright is invoked.

```

## Summary

- The `SmartRouter` in [`src/fetch/router.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/router.ts) orchestrates a **three-tier fetch strategy**: HTTP → TLS → Playwright, escalating only when necessary.
- **Anti-bot detection** relies on `isAntiBotSignal` and `isChallengeShell` from [`src/fetch/tls-tier.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/tls-tier.ts) to identify 403/429/503 status codes and interstitial HTML patterns even in 2xx responses.
- **SPA handling** combines a static `KNOWN_SPA_DOMAINS` list with dynamic `DomainStats` learning to minimize unnecessary browser launches while ensuring JavaScript-rendered content is captured.
- **Stealth mode** (`mode: 'stealth'`) defers expensive Playwright instantiation until content analysis confirms thin or challenged responses, optimizing resource usage.
- Failed challenges are converted to structured errors via `guardChallengeShell` (lines 886-896) rather than returning raw HTML interstitials.

## Frequently Asked Questions

### What triggers the TLS tier in Wigolo's fetch router?

The TLS tier activates when `getConfig().tlsTier` is set to `'on'` or when the target domain appears in the anti-bot allow-list verified by `isAntiBotTlsDomain` (lines 106-108 in [`src/fetch/router.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/fetch/router.ts)). When triggered, the router attempts TLS impersonation via `tryTlsTier` before falling back to Playwright for challenge-heavy sites.

### How does Wigolo avoid launching Playwright for every request?

The router uses the `KNOWN_SPA_DOMAINS` array to short-circuit HTTP attempts for known JavaScript frameworks, while per-host `DomainStats` tracked in memory allow dynamic downgrading from Playwright to HTTP when static content proves sufficient. Additionally, `mode: 'stealth'` performs static fetches first and only escalates when `shouldEscalate` or `isChallengeShell` returns true.

### What happens when Wigolo encounters a Cloudflare challenge page?

When `isChallengeShell` detects Cloudflare's "Just a moment" interstitial (even with a 200 status code), the router first attempts TLS impersonation if enabled via the configuration. If TLS fails to bypass the challenge, it escalates to `browserFetch`, which uses `guardChallengeShell` (lines 886-896) to return a structured `blocked_by_challenge` error rather than the raw HTML.

### Can I force Playwright for specific domains while keeping HTTP as default for others?

Yes. You can pass `renderJs: 'always'` in your fetch options for specific requests, or manipulate the `DomainStats` preference system to set `preferPlaywright = true` for particular domains. The router respects per-domain statistics over global defaults, allowing fine-grained control over which tier handles each host.