How archive.today Fallback Handles CAPTCHAs in qiaomu-anything-to-notebooklm
The archive.today fallback detects CAPTCHA challenges using pattern matching against security keywords and signals the need for manual intervention by exiting with code 75, enabling upstream callers to prompt for human verification rather than returning incomplete data.
The qiaomu-anything-to-notebooklm repository implements a robust six-level cascade to retrieve pay-walled web content for processing into NotebookLM-compatible formats. When direct fetching methods fail, the system activates the archive.today fallback CAPTCHA handling mechanism, which distinguishes between valid archived snapshots and anti-bot security challenges. This architecture ensures automated pipelines degrade gracefully rather than silently returning placeholder content or error pages.
The Six-Level Retrieval Cascade
Position of the archive.today Fallback
Level 4 of the retrieval cascade specifically targets archive.today (archive.ph) snapshots when earlier bypass attempts fail. The script constructs the snapshot request using https://archive.today/newest/$URL, executing a timed HTTP GET with a 20-second timeout and a desktop User-Agent header to mimic browser traffic.
ARCHIVE_URL="https://archive.today/newest/$URL"
ARCHIVE_OUT=$(_curl -sL "$ARCHIVE_URL" \
-H "User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36" \
--max-time 20 2>/dev/null || true)
CAPTCHA Detection and Content Validation
Pattern-Based Challenge Detection
After fetching the snapshot, the script invokes _is_captcha_page to scan the HTML response for security challenge indicators. According to the implementation in scripts/fetch_url.sh, this function performs case-insensitive pattern matching against strings including "security check", "captcha", "recaptcha", and "cloudflare.*challenge" to identify bot detection pages.
Content Sanity Verification
Simultaneously, the _has_content function validates that the response contains sufficient line and character counts to qualify as a substantive article rather than an error or placeholder page. Only content passing both the CAPTCHA screen and the sanity check proceeds to extraction via _html_to_text.
Signaling Manual Intervention with Exit Code 75
When _is_captcha_page identifies a security challenge, the script aborts the automated pipeline with a specific diagnostic protocol defined in scripts/fetch_url.sh. Instead of returning partial data, the system writes ARCHIVE_CAPTCHA:$ARCHIVE_URL to standard error, prints instructions requesting browser-based verification, and terminates with exit code 75.
echo "ARCHIVE_CAPTCHA:$ARCHIVE_URL" >&2
echo "⚠️ archive.ph needs human verification." >&2
echo " The caller should open this URL in a browser for the user to solve the CAPTCHA," >&2
echo " then retry this script." >&2
exit 75
This exit code is recognized by the surrounding Claude skill as a "human-verification required" signal. The caller extracts the archive URL from stderr, presents it to the user for manual solving, and re-invokes the script once access is granted.
Practical Implementation Examples
Bash Wrapper for CAPTCHA Handling
The following shell script demonstrates how to invoke the fetch utility and handle the CAPTCHA signal appropriately:
#!/usr/bin/env bash
URL="https://example.com/paywalled-article"
./scripts/fetch_url.sh "$URL"
RET=$?
if [[ $RET -eq 75 ]]; then
echo "⚠️ Archive.today requires CAPTCHA verification."
CAP_URL=$(./scripts/fetch_url.sh "$URL" 2>&1 | grep '^ARCHIVE_CAPTCHA' | cut -d: -f2-)
echo "Open the following URL in a browser to solve the CAPTCHA:"
echo "$CAP_URL"
echo "After solving, rerun the script."
fi
Python Integration
For Python-based automation, this wrapper captures the special exit code and raises a descriptive exception:
import subprocess
import sys
def fetch_with_archive(url: str) -> str:
proc = subprocess.run(
["./scripts/fetch_url.sh", url],
capture_output=True,
text=True,
)
if proc.returncode == 75:
cap_url = next(
line.split(":", 1)[1] for line in proc.stderr.splitlines()
if line.startswith("ARCHIVE_CAPTCHA:")
)
raise RuntimeError(
f"CAPTCHA required – please open {cap_url} in a browser and retry."
)
proc.check_returncode()
return proc.stdout
if __name__ == "__main__":
try:
print(fetch_with_archive(sys.argv[1]))
except RuntimeError as e:
print(e, file=sys.stderr)
sys.exit(1)
Summary
- The archive.today fallback operates at Level 4 of the six-level retrieval cascade implemented in
scripts/fetch_url.sh. - CAPTCHA detection relies on
_is_captcha_pagescanning HTML for keywords including "recaptcha" and "cloudflare.*challenge". - Valid content must pass both challenge detection and content sanity verification via
_has_contentbefore text extraction proceeds. - Upon detecting a CAPTCHA, the system emits
ARCHIVE_CAPTCHA:<url>to stderr and exits with code 75, enabling upstream callers to initiate manual browser-based verification. - This architecture prevents data corruption from incomplete fetches and provides clear, actionable feedback when automated retrieval encounters security barriers.
Frequently Asked Questions
What exit code indicates a CAPTCHA challenge from archive.today?
The script exits with code 75 when archive.today returns a CAPTCHA page. This specific code is recognized by the Claude skill interface as indicating that human verification is required, allowing calling applications to distinguish between permanent fetch failures and temporary blocks solvable via manual intervention.
How does the script distinguish between valid content and a CAPTCHA page?
The _is_captcha_page function in scripts/fetch_url.sh performs case-insensitive pattern matching against the HTTP response body, searching for strings such as "security check", "captcha", "recaptcha", and "cloudflare.*challenge". Simultaneously, _has_content verifies the response contains sufficient text volume to represent a genuine article rather than a placeholder or error page.
Will the fallback return partial content if a CAPTCHA is detected?
No. When the script detects a CAPTCHA, it explicitly prevents content extraction and text conversion. It returns exit code 75 and diagnostic messages to stderr rather than risking the return of incomplete data, security check pages, or misleading placeholders that could contaminate the final notebook output.
How should automation pipelines handle exit code 75?
Capture the specific archive URL from the stderr output where the script prints ARCHIVE_CAPTCHA:<url>, present this URL to the user for manual browser-based verification, and re-execute the fetch operation once the CAPTCHA is solved. Both the Bash and Python examples above demonstrate parsing the stderr stream to extract the verification URL and pause the pipeline appropriately.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →