How the AI Job Search Framework Handles HTTP 403 Errors When Fetching Job Postings

A HTTP 403 error during robots.txt fetching results in immediate abort with no retry, as the framework treats non-200 responses as unconfirmed permission.

When scraping job postings at scale, encountering HTTP 403 Forbidden errors is inevitable. The MadsLorentzen/ai-job-search framework implements a deliberately cautious approach to handling these errors, prioritizing ethical crawling over aggressive retries. Rather than attempting to bypass restrictions, the framework checks robots.txt compliance first and treats any ambiguous response—including 403 status codes—as a signal to stop.

How the Framework Detects and Handles HTTP 403 Errors

The framework's error handling logic lives in tools/robots_check.py, which serves as a gatekeeper before any actual job posting fetch occurs. This design ensures that HTTP 403 errors encountered during the permission-checking phase automatically prevent retries.

The robots.txt Check as First-Line Defense

Before fetching any job posting URL, the framework queries the host's robots.txt file using two user-agents:

  • Claude-User — the framework's internal identity
  • BROWSER — a generic browser string for fallback comparison

The _fetch() function uses curl rather than urllib for reliability:


# tools/robots_check.py, lines 30-41

r = subprocess.run(
    ['curl', '-sS', '-L', '--max-redirs', '5', '--max-time', '12',
     '-A', ua, '-H', 'Accept: text/plain,*/*',
     '-w', '\n%{http_code}', '--', url],
    capture_output=True, text=True, timeout=20)

HTTP Status Code Interpretation

The framework classifies responses into three categories:

Status Interpretation Action
404 No robots.txt published Allowed — proceed with retry (exit 0)
200 robots.txt present Parse rules and apply user-agent matching
403 or any other Permission unconfirmed Blocked — no retry, escalate (exit 1)

The critical logic appears in lines 21-30 of tools/robots_check.py:


# tools/robots_check.py, lines 21-30

if code == 404:                      # no policy → permission granted

    return 0, 'ALLOWED - no robots.txt published'
if code == 200:
    ...  # parse and apply rules

# any other code (e.g., 403) falls through

return 1, f'UNCONFIRMED ({last}) - do not retry, go to step 3'

This means a HTTP 403 error when fetching robots.txt immediately terminates the retry attempt with exit code 1.

Rule Parsing When 200 is Returned

If robots.txt returns HTTP 200, the parser extracts Allow/Disallow directives per user-agent. The matching algorithm uses longest-match priority with a cautious default: on a tie, Disallow overrides Allow.


# tools/robots_check.py, lines 99-107

for is_allow, pat in rules:
    n = _match(pat, path)
    if n > best_len or (n == best_len and n >= 0 and not is_allow):
        best_len, best_allow = n, is_allow  # Disallow wins on tie

Practical Implementation Examples

Command-Line Usage

Check a URL before fetching from the repository root:

python3 -m tools.robots_check https://example.com/job/12345

Exit codes indicate action:

  • 0 — retry may proceed
  • 1 — do not retry (403, disallowed, or other error)
  • 2 — usage error

Programmatic Integration

import subprocess, sys
from tools import robots_check

def may_retry(url: str) -> bool:
    rc, _ = robots_check.gate(url)      # returns (0|1, message)

    return rc == 0

# Example:

if may_retry("https://jobs.example.com/software-engineer"):
    # safe to issue the actual job-posting request

    ...
else:
    print("Access denied or uncertain – aborting fetch.")

Why HTTP 403 Errors Trigger Immediate Abort

The framework's behavior stems from RFC 9309 compliance interpreted conservatively. A HTTP 403 response when fetching robots.txt indicates one of several conditions:

  • The server actively blocks automated requests
  • The robots.txt endpoint requires authentication
  • Rate limiting or IP-based restrictions apply

In all cases, the framework cannot confirm permission to crawl, so it defaults to denial rather than risk violating terms of service.

Key Files and Components

File Role in HTTP 403 Handling
tools/robots_check.py Core logic: fetches robots.txt, interprets HTTP codes including 403, decides retry permission
tests/test_robots_check.py Validates correct handling of 403, 404, and other status codes
_fetch (internal function) curl wrapper providing reliable HTTP with timeout and error handling
README.md Documents the retry policy and ethical crawling workflow

Summary

  • HTTP 403 errors abort retries immediately — no automatic bypass or retry logic exists
  • robots.txt checking precedes all job posting fetches as mandatory gate
  • Exit code 1 signals unconfirmed permission for 403 responses and other non-200/404 statuses
  • Cautious default philosophy prioritizes ethical compliance over crawl completeness
  • tools/robots_check.py contains all relevant logic for HTTP 403 handling

Frequently Asked Questions

What happens if a job site returns HTTP 403 for the actual posting but not robots.txt?

The framework never reaches the posting fetch if robots_check.py encounters ambiguity. A 403 on the posting itself would occur outside this module's scope—the framework separates permission checking from content fetching. If robots.txt returns 200 and allows the path, other HTTP errors during posting retrieval are handled by separate retry logic with exponential backoff.

Can the framework be configured to retry on HTTP 403 errors?

No. The exit code 1 from tools/robots_check.py is hardcoded for all non-200/404 responses. This design choice reflects the repository's ethical crawling principles. Modification would require editing the fallback return statement in lines 29-30 of tools/robots_check.py.

How does the framework distinguish between robots.txt 403 and posting 403?

The robots_check.py module only fetches https://<host>/robots.txt. A 403 at this specific endpoint triggers the unconfirmed permission path. The actual job posting URL is never requested within this module—subsequent fetching occurs only after gate() returns exit code 0.

Why use curl instead of Python's urllib for robots.txt fetching?

The _fetch() implementation in tools/robots_check.py uses curl for robust timeout handling, redirect following (--max-redirs 5), and explicit user-agent control. This avoids Python's urllib complexities with SSL certificates and provides more predictable error codes for downstream HTTP 403 detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →