How Wigolo Per-Domain Fetch Learning Works: Inspection and Reset Guide

Wigolo automatically learns optimal fetch strategies for each domain by persisting TLS success counters, anti-bot clearance cookies, and backoff timers in a local SQLite database, which you can inspect or clear using the wigolo tune CLI command.

Wigolo's fetch router adapts its behavior per domain based on historical success and failure patterns. This learning system stores routing preferences, clearance tokens, and failure statistics in a SQLite table named domain_routing, enabling the tool to optimize between lightweight TLS impersonation and full browser automation. Understanding how this mechanism works—and how to reset it when troubleshooting—is essential for maintaining reliable data extraction pipelines.

The SQLite Storage Architecture

All per-domain learning data lives in a SQLite database located at <dataDir>/wigolo.db (default: ~/.wigolo/wigolo.db). The schema is defined in migration 008-antibot-clearance.sql and maintained through src/cache/store.ts.

-- src/cache/migrations/008-antibot-clearance.sql
CREATE TABLE domain_routing (
  domain TEXT PRIMARY KEY,
  prefer_playwright INTEGER,
  http_failures INTEGER,
  prefer_tls_impersonation INTEGER,
  tls_success_count INTEGER,
  backoff_until TEXT,
  last_403_at TEXT,
  cf_clearance TEXT,
  clearance_ua TEXT,
  clearance_tier TEXT,
  clearance_expires_at TEXT,
  last_updated TEXT
);

This table tracks five key aspects of wigolo's per-domain fetch learning:

  • TLS preference: Whether to prefer TLS impersonation over browser automation
  • Browser preference: Whether Playwright should be the default tier
  • Anti-bot clearance: Cloudflare (or similar) clearance cookies with expiry and user agent
  • Backoff timing: Cool-down periods after repeated 403 responses
  • Statistics: Counters for TLS successes and HTTP failures

How Learning Mechanisms Record Domain Behavior

Tracking TLS Impersonation Success

When the TLS-impersonation tier successfully fetches a URL, src/fetch/router.ts calls recordTlsImpersonationSuccess() to increment a success counter. Once the counter reaches the configured threshold (defined by tlsSuccessThreshold), the system atomically sets prefer_tls_impersonation to true for that domain.

// src/cache/store.ts
export function recordTlsImpersonationSuccess(domain: string, threshold: number): DomainRoutingRow | null {
  const db = getDatabase();
  const preferOnInsert = threshold <= 1 ? 1 : 0;
  db.prepare(`
    INSERT INTO domain_routing (domain, prefer_playwright, http_failures,
                               prefer_tls_impersonation, tls_success_count, last_updated)
    VALUES (?, 0, 0, ?, 1, datetime('now'))
    ON CONFLICT(domain) DO UPDATE SET
      tls_success_count = tls_success_count + 1,
      last_updated = datetime('now'),
      prefer_tls_impersonation = CASE
        WHEN tls_success_count + 1 >= ?
        THEN 1
        ELSE prefer_tls_impersonation
      END
  `).run(domain, preferOnInsert, threshold);
  return getDomainRouting(domain);
}

The router invokes this function after successful TLS fetches:

// src/fetch/router.ts
recordSuccess(domain) {
  try {
    recordTlsImpersonationSuccess(domain, getConfig().tlsSuccessThreshold);
  } catch { /* best-effort */ }
}

Storing Anti-Bot Clearances

When the browser tier solves a challenge (such as Cloudflare), Wigolo caches the clearance cookie, user agent, and expiry to reuse for subsequent requests:

// src/cache/store.ts
export function recordDomainClearance(host: string, clearance: DomainClearance): void {
  const db = getDatabase();
  db.prepare(`
    INSERT INTO domain_routing (domain, prefer_playwright, http_failures,
                               cf_clearance, clearance_ua, clearance_tier,
                               clearance_expires_at, last_updated)
    VALUES (?, 0, 0, ?, ?, ?, ?, datetime('now'))
    ON CONFLICT(domain) DO UPDATE SET
      cf_clearance = excluded.cf_clearance,
      clearance_ua = excluded.clearance_ua,
      clearance_tier = excluded.clearance_tier,
      clearance_expires_at = excluded.clearance_expires_at,
      last_updated = datetime('now')
  `).run(host, clearance.cookie, clearance.ua, clearance.tier, clearance.expiresAt);
}

Retrieval hides sensitive cookie values but reports presence and expiry:

// src/cache/store.ts
export function getDomainClearance(host: string): DomainClearance | null {
  const row = db.prepare(`
    SELECT cf_clearance, clearance_ua, clearance_tier, clearance_expires_at
    FROM domain_routing WHERE domain = ? LIMIT 1`).get(host);
  if (!row || row.cf_clearance == null) return null;
  return {
    cookie: row.cf_clearance,
    ua: row.clearance_ua ?? '',
    tier: row.clearance_tier ?? '',
    expiresAt: row.clearance_expires_at ?? '',
  };
}

Managing Backoff After Blocks

Repeated 403 responses trigger a cool-down period recorded via recordBackoff():

// src/cache/store.ts
export function recordBackoff(host: string, untilEpochMs: number): void {
  const db = getDatabase();
  db.prepare(`
    INSERT INTO domain_routing (domain, prefer_playwright, http_failures,
                               backoff_until, last_403_at, last_updated)
    VALUES (?, 0, 0, ?, datetime('now'), datetime('now'))
    ON CONFLICT(domain) DO UPDATE SET
      backoff_until = excluded.backoff_until,
      last_403_at = datetime('now'),
      last_updated = datetime('now')
  `).run(host, String(untilEpochMs));
}

The fetch ladder checks getBackoff(host) before attempting requests to respect polite delays.

Inspecting Learned Routing with wigolo tune

The wigolo tune command provides read-only access to the learning database through src/cli/tune.ts. The CLI intentionally omits raw cookie values and user agents from output for security, showing only whether clearances are present.

Listing All Domains

wigolo tune list

This executes listDomainRouting() in src/cache/store.ts, which returns a summary projection:

// src/cache/store.ts
export function listDomainRouting(): DomainRoutingSummary[] {
  const rows = db.prepare(`
    SELECT domain, prefer_playwright, prefer_tls_impersonation, tls_success_count,
           http_failures, backoff_until, last_403_at, cf_clearance,
           clearance_expires_at
    FROM domain_routing ORDER BY domain`).all();
  return rows.map(row => ({
    domain: row.domain,
    preferBrowser: (row.prefer_playwright ?? 0) === 1,
    preferTlsImpersonation: (row.prefer_tls_impersonation ?? 0) === 1,
    tlsSuccessCount: row.tls_success_count ?? 0,
    httpFailures: row.http_failures ?? 0,
    backoffUntil: row.backoff_until ?? undefined,
    last403At: row.last_403_at ?? undefined,
    clearancePresent: row.cf_clearance != null,
    clearanceExpiresAt: row.clearance_expires_at ?? undefined,
  }));
}

Viewing Specific Domain Details

wigolo tune show example.com --json

Sample output showing learned preferences:

{
  "domain": "example.com",
  "preferBrowser": false,
  "preferTlsImpersonation": true,
  "tlsSuccessCount": 3,
  "httpFailures": 0,
  "clearancePresent": true,
  "clearanceExpiresAt": "2026-12-01 12:34:56"
}

Clearing or Resetting Domain Learning

When domains change their protection mechanisms or you need to troubleshoot fetch failures, you can reset the learned state without deleting the database.

Reset Single Domain

wigolo tune reset example.com

This calls resetDomainRouting() in src/cache/store.ts:

// src/cache/store.ts
export function resetDomainRouting(host: string): number {
  return db.prepare(`UPDATE domain_routing SET ${RESET_ROUTING_COLUMNS} WHERE domain = ?`)
    .run(host).changes;
}

Reset All Domains

wigolo tune reset --all

This invokes resetAllDomainRouting():

// src/cache/store.ts
export function resetAllDomainRouting(): number {
  return db.prepare(`UPDATE domain_routing SET ${RESET_ROUTING_COLUMNS}`).run().changes;
}

Both functions preserve the domain row in the table but nullify all learned columns, allowing fresh learning to begin immediately.

Summary

Wigolo's per-domain fetch learning system automatically optimizes request strategies through persistent SQLite storage:

  • TLS tracking increments tls_success_count and sets prefer_tls_impersonation when domains consistently accept TLS impersonation
  • Clearance caching stores solved anti-bot cookies with expiry timestamps to minimize repeated browser automation
  • Backoff management records cool-down periods after 403 responses to prevent rapid retry loops
  • Inspection tools via wigolo tune list and show reveal current preferences without exposing sensitive cookie values
  • Reset capabilities allow clearing learned state per-domain or globally using wigolo tune reset

Frequently Asked Questions

How do I check if Wigolo has learned to use TLS impersonation for a specific domain?

Run wigolo tune show <domain> --json and check the preferTlsImpersonation field. If true, the fetch router in src/fetch/router.ts will prefer the TLS tier over browser automation for that domain based on previous successful fetches recorded in tls_success_count.

Where is the per-domain learning data physically stored?

The data resides in a SQLite database at <dataDir>/wigolo.db, defaulting to ~/.wigolo/wigolo.db. The domain_routing table defined in src/cache/migrations/008-antibot-clearance.sql contains all counters, preferences, and clearance cookies managed by src/cache/store.ts.

The listDomainRouting() and related functions intentionally exclude raw cf_clearance values and user agents from CLI output to prevent accidental exposure of session tokens in terminal logs or shared environments. The JSON output only indicates clearancePresent: true with expiry timing.

When should I reset domain learning data?

Reset specific domains using wigolo tune reset <domain> when sites change their protection mechanisms, when you suspect stale clearance cookies are causing blocks, or when debugging fetch failures after TLS or anti-bot updates. Use wigolo tune reset --all only when you want to force fresh learning across all previously visited domains.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →