How the AI Job Search Framework Handles HTTP 403 Errors When Fetching Job Postings
A HTTP 403 error during robots.txt fetching results in immediate abort with no retry, as the framework treats non-200 responses as unconfirmed permission.
When scraping job postings at scale, encountering HTTP 403 Forbidden errors is inevitable. The MadsLorentzen/ai-job-search framework implements a deliberately cautious approach to handling these errors, prioritizing ethical crawling over aggressive retries. Rather than attempting to bypass restrictions, the framework checks robots.txt compliance first and treats any ambiguous response—including 403 status codes—as a signal to stop.
How the Framework Detects and Handles HTTP 403 Errors
The framework's error handling logic lives in tools/robots_check.py, which serves as a gatekeeper before any actual job posting fetch occurs. This design ensures that HTTP 403 errors encountered during the permission-checking phase automatically prevent retries.
The robots.txt Check as First-Line Defense
Before fetching any job posting URL, the framework queries the host's robots.txt file using two user-agents:
Claude-User— the framework's internal identityBROWSER— a generic browser string for fallback comparison
The _fetch() function uses curl rather than urllib for reliability:
# tools/robots_check.py, lines 30-41
r = subprocess.run(
['curl', '-sS', '-L', '--max-redirs', '5', '--max-time', '12',
'-A', ua, '-H', 'Accept: text/plain,*/*',
'-w', '\n%{http_code}', '--', url],
capture_output=True, text=True, timeout=20)
HTTP Status Code Interpretation
The framework classifies responses into three categories:
| Status | Interpretation | Action |
|---|---|---|
| 404 | No robots.txt published | Allowed — proceed with retry (exit 0) |
| 200 | robots.txt present | Parse rules and apply user-agent matching |
| 403 or any other | Permission unconfirmed | Blocked — no retry, escalate (exit 1) |
The critical logic appears in lines 21-30 of tools/robots_check.py:
# tools/robots_check.py, lines 21-30
if code == 404: # no policy → permission granted
return 0, 'ALLOWED - no robots.txt published'
if code == 200:
... # parse and apply rules
# any other code (e.g., 403) falls through
return 1, f'UNCONFIRMED ({last}) - do not retry, go to step 3'
This means a HTTP 403 error when fetching robots.txt immediately terminates the retry attempt with exit code 1.
Rule Parsing When 200 is Returned
If robots.txt returns HTTP 200, the parser extracts Allow/Disallow directives per user-agent. The matching algorithm uses longest-match priority with a cautious default: on a tie, Disallow overrides Allow.
# tools/robots_check.py, lines 99-107
for is_allow, pat in rules:
n = _match(pat, path)
if n > best_len or (n == best_len and n >= 0 and not is_allow):
best_len, best_allow = n, is_allow # Disallow wins on tie
Practical Implementation Examples
Command-Line Usage
Check a URL before fetching from the repository root:
python3 -m tools.robots_check https://example.com/job/12345
Exit codes indicate action:
0— retry may proceed1— do not retry (403, disallowed, or other error)2— usage error
Programmatic Integration
import subprocess, sys
from tools import robots_check
def may_retry(url: str) -> bool:
rc, _ = robots_check.gate(url) # returns (0|1, message)
return rc == 0
# Example:
if may_retry("https://jobs.example.com/software-engineer"):
# safe to issue the actual job-posting request
...
else:
print("Access denied or uncertain – aborting fetch.")
Why HTTP 403 Errors Trigger Immediate Abort
The framework's behavior stems from RFC 9309 compliance interpreted conservatively. A HTTP 403 response when fetching robots.txt indicates one of several conditions:
- The server actively blocks automated requests
- The
robots.txtendpoint requires authentication - Rate limiting or IP-based restrictions apply
In all cases, the framework cannot confirm permission to crawl, so it defaults to denial rather than risk violating terms of service.
Key Files and Components
| File | Role in HTTP 403 Handling |
|---|---|
tools/robots_check.py |
Core logic: fetches robots.txt, interprets HTTP codes including 403, decides retry permission |
tests/test_robots_check.py |
Validates correct handling of 403, 404, and other status codes |
_fetch (internal function) |
curl wrapper providing reliable HTTP with timeout and error handling |
README.md |
Documents the retry policy and ethical crawling workflow |
Summary
- HTTP 403 errors abort retries immediately — no automatic bypass or retry logic exists
- robots.txt checking precedes all job posting fetches as mandatory gate
- Exit code 1 signals unconfirmed permission for 403 responses and other non-200/404 statuses
- Cautious default philosophy prioritizes ethical compliance over crawl completeness
tools/robots_check.pycontains all relevant logic for HTTP 403 handling
Frequently Asked Questions
What happens if a job site returns HTTP 403 for the actual posting but not robots.txt?
The framework never reaches the posting fetch if robots_check.py encounters ambiguity. A 403 on the posting itself would occur outside this module's scope—the framework separates permission checking from content fetching. If robots.txt returns 200 and allows the path, other HTTP errors during posting retrieval are handled by separate retry logic with exponential backoff.
Can the framework be configured to retry on HTTP 403 errors?
No. The exit code 1 from tools/robots_check.py is hardcoded for all non-200/404 responses. This design choice reflects the repository's ethical crawling principles. Modification would require editing the fallback return statement in lines 29-30 of tools/robots_check.py.
How does the framework distinguish between robots.txt 403 and posting 403?
The robots_check.py module only fetches https://<host>/robots.txt. A 403 at this specific endpoint triggers the unconfirmed permission path. The actual job posting URL is never requested within this module—subsequent fetching occurs only after gate() returns exit code 0.
Why use curl instead of Python's urllib for robots.txt fetching?
The _fetch() implementation in tools/robots_check.py uses curl for robust timeout handling, redirect following (--max-redirs 5), and explicit user-agent control. This avoids Python's urllib complexities with SSL certificates and provides more predictable error codes for downstream HTTP 403 detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →