# How the AI Job Search Framework Handles HTTP 403 Errors When Fetching Job Postings

> Learn how the AI Job Search framework handles HTTP 403 errors. Discover why it aborts immediately and avoids retries when fetching job postings, ensuring permission confirmation.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-08-30

---

**A HTTP 403 error during robots.txt fetching results in immediate abort with no retry, as the framework treats non-200 responses as unconfirmed permission.**

When scraping job postings at scale, encountering HTTP 403 Forbidden errors is inevitable. The `MadsLorentzen/ai-job-search` framework implements a deliberately cautious approach to handling these errors, prioritizing ethical crawling over aggressive retries. Rather than attempting to bypass restrictions, the framework checks [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) compliance first and treats any ambiguous response—including 403 status codes—as a signal to stop.

## How the Framework Detects and Handles HTTP 403 Errors

The framework's error handling logic lives in **[`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py)**, which serves as a gatekeeper before any actual job posting fetch occurs. This design ensures that HTTP 403 errors encountered during the permission-checking phase automatically prevent retries.

### The robots.txt Check as First-Line Defense

Before fetching any job posting URL, the framework queries the host's [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) file using two user-agents:

- **`Claude-User`** — the framework's internal identity
- **`BROWSER`** — a generic browser string for fallback comparison

The `_fetch()` function uses `curl` rather than `urllib` for reliability:

```python

# tools/robots_check.py, lines 30-41

r = subprocess.run(
    ['curl', '-sS', '-L', '--max-redirs', '5', '--max-time', '12',
     '-A', ua, '-H', 'Accept: text/plain,*/*',
     '-w', '\n%{http_code}', '--', url],
    capture_output=True, text=True, timeout=20)

```

### HTTP Status Code Interpretation

The framework classifies responses into three categories:

| Status | Interpretation | Action |
|--------|---------------|--------|
| **404** | No robots.txt published | **Allowed** — proceed with retry (exit 0) |
| **200** | robots.txt present | Parse rules and apply user-agent matching |
| **403 or any other** | Permission unconfirmed | **Blocked** — no retry, escalate (exit 1) |

The critical logic appears in lines 21-30 of [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py):

```python

# tools/robots_check.py, lines 21-30

if code == 404:                      # no policy → permission granted

    return 0, 'ALLOWED - no robots.txt published'
if code == 200:
    ...  # parse and apply rules

# any other code (e.g., 403) falls through

return 1, f'UNCONFIRMED ({last}) - do not retry, go to step 3'

```

This means a **HTTP 403 error when fetching robots.txt** immediately terminates the retry attempt with exit code 1.

### Rule Parsing When 200 is Returned

If [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) returns HTTP 200, the parser extracts Allow/Disallow directives per user-agent. The matching algorithm uses longest-match priority with a **cautious default**: on a tie, Disallow overrides Allow.

```python

# tools/robots_check.py, lines 99-107

for is_allow, pat in rules:
    n = _match(pat, path)
    if n > best_len or (n == best_len and n >= 0 and not is_allow):
        best_len, best_allow = n, is_allow  # Disallow wins on tie

```

## Practical Implementation Examples

### Command-Line Usage

Check a URL before fetching from the repository root:

```bash
python3 -m tools.robots_check https://example.com/job/12345

```

Exit codes indicate action:
- `0` — retry may proceed
- `1` — do not retry (403, disallowed, or other error)
- `2` — usage error

### Programmatic Integration

```python
import subprocess, sys
from tools import robots_check

def may_retry(url: str) -> bool:
    rc, _ = robots_check.gate(url)      # returns (0|1, message)

    return rc == 0

# Example:

if may_retry("https://jobs.example.com/software-engineer"):
    # safe to issue the actual job-posting request

    ...
else:
    print("Access denied or uncertain – aborting fetch.")

```

## Why HTTP 403 Errors Trigger Immediate Abort

The framework's behavior stems from RFC 9309 compliance interpreted conservatively. A HTTP 403 response when fetching [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) indicates one of several conditions:

- The server actively blocks automated requests
- The [`robots.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots.txt) endpoint requires authentication
- Rate limiting or IP-based restrictions apply

In all cases, the framework cannot confirm permission to crawl, so it **defaults to denial** rather than risk violating terms of service.

## Key Files and Components

| File | Role in HTTP 403 Handling |
|------|--------------------------|
| [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py) | Core logic: fetches robots.txt, interprets HTTP codes including 403, decides retry permission |
| [`tests/test_robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_robots_check.py) | Validates correct handling of 403, 404, and other status codes |
| `_fetch` (internal function) | `curl` wrapper providing reliable HTTP with timeout and error handling |
| [`README.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/README.md) | Documents the retry policy and ethical crawling workflow |

## Summary

- **HTTP 403 errors abort retries immediately** — no automatic bypass or retry logic exists
- **robots.txt checking precedes all job posting fetches** as mandatory gate
- **Exit code 1 signals unconfirmed permission** for 403 responses and other non-200/404 statuses
- **Cautious default philosophy** prioritizes ethical compliance over crawl completeness
- **[`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py)** contains all relevant logic for HTTP 403 handling

## Frequently Asked Questions

### What happens if a job site returns HTTP 403 for the actual posting but not robots.txt?

The framework never reaches the posting fetch if [`robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots_check.py) encounters ambiguity. A 403 on the posting itself would occur outside this module's scope—the framework separates permission checking from content fetching. If robots.txt returns 200 and allows the path, other HTTP errors during posting retrieval are handled by separate retry logic with exponential backoff.

### Can the framework be configured to retry on HTTP 403 errors?

No. The exit code 1 from [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py) is hardcoded for all non-200/404 responses. This design choice reflects the repository's ethical crawling principles. Modification would require editing the fallback return statement in lines 29-30 of [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py).

### How does the framework distinguish between robots.txt 403 and posting 403?

The [`robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/robots_check.py) module only fetches `https://<host>/robots.txt`. A 403 at this specific endpoint triggers the unconfirmed permission path. The actual job posting URL is never requested within this module—subsequent fetching occurs only after `gate()` returns exit code 0.

### Why use curl instead of Python's urllib for robots.txt fetching?

The `_fetch()` implementation in [`tools/robots_check.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/robots_check.py) uses `curl` for robust timeout handling, redirect following (`--max-redirs 5`), and explicit user-agent control. This avoids Python's urllib complexities with SSL certificates and provides more predictable error codes for downstream HTTP 403 detection.