# How archive.today Fallback Handles CAPTCHAs in qiaomu-anything-to-notebooklm

> Discover how archive.today fallback handles CAPTCHAs using pattern matching. Learn how it enables human verification for incomplete data in joeseesun/qiaomu-anything-to-notebooklm.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: how-to-guide
- Published: 2026-05-16

---

**The archive.today fallback detects CAPTCHA challenges using pattern matching against security keywords and signals the need for manual intervention by exiting with code 75, enabling upstream callers to prompt for human verification rather than returning incomplete data.**

The qiaomu-anything-to-notebooklm repository implements a robust six-level cascade to retrieve pay-walled web content for processing into NotebookLM-compatible formats. When direct fetching methods fail, the system activates the **archive.today fallback CAPTCHA handling** mechanism, which distinguishes between valid archived snapshots and anti-bot security challenges. This architecture ensures automated pipelines degrade gracefully rather than silently returning placeholder content or error pages.

## The Six-Level Retrieval Cascade

### Position of the archive.today Fallback

Level 4 of the retrieval cascade specifically targets archive.today (archive.ph) snapshots when earlier bypass attempts fail. The script constructs the snapshot request using `https://archive.today/newest/$URL`, executing a timed HTTP GET with a 20-second timeout and a desktop User-Agent header to mimic browser traffic.

```bash
ARCHIVE_URL="https://archive.today/newest/$URL"
ARCHIVE_OUT=$(_curl -sL "$ARCHIVE_URL" \
  -H "User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36" \
  --max-time 20 2>/dev/null || true)

```

## CAPTCHA Detection and Content Validation

### Pattern-Based Challenge Detection

After fetching the snapshot, the script invokes `_is_captcha_page` to scan the HTML response for security challenge indicators. According to the implementation in [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh), this function performs case-insensitive pattern matching against strings including "security check", "captcha", "recaptcha", and "cloudflare.*challenge" to identify bot detection pages.

### Content Sanity Verification

Simultaneously, the `_has_content` function validates that the response contains sufficient line and character counts to qualify as a substantive article rather than an error or placeholder page. Only content passing both the CAPTCHA screen and the sanity check proceeds to extraction via `_html_to_text`.

## Signaling Manual Intervention with Exit Code 75

When `_is_captcha_page` identifies a security challenge, the script aborts the automated pipeline with a specific diagnostic protocol defined in [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh). Instead of returning partial data, the system writes `ARCHIVE_CAPTCHA:$ARCHIVE_URL` to standard error, prints instructions requesting browser-based verification, and terminates with **exit code 75**.

```bash
echo "ARCHIVE_CAPTCHA:$ARCHIVE_URL" >&2
echo "⚠️  archive.ph needs human verification." >&2
echo "   The caller should open this URL in a browser for the user to solve the CAPTCHA," >&2
echo "   then retry this script." >&2
exit 75

```

This exit code is recognized by the surrounding Claude skill as a "human-verification required" signal. The caller extracts the archive URL from stderr, presents it to the user for manual solving, and re-invokes the script once access is granted.

## Practical Implementation Examples

### Bash Wrapper for CAPTCHA Handling

The following shell script demonstrates how to invoke the fetch utility and handle the CAPTCHA signal appropriately:

```bash
#!/usr/bin/env bash
URL="https://example.com/paywalled-article"
./scripts/fetch_url.sh "$URL"
RET=$?

if [[ $RET -eq 75 ]]; then
  echo "⚠️ Archive.today requires CAPTCHA verification."
  CAP_URL=$(./scripts/fetch_url.sh "$URL" 2>&1 | grep '^ARCHIVE_CAPTCHA' | cut -d: -f2-)
  echo "Open the following URL in a browser to solve the CAPTCHA:"
  echo "$CAP_URL"
  echo "After solving, rerun the script."
fi

```

### Python Integration

For Python-based automation, this wrapper captures the special exit code and raises a descriptive exception:

```python
import subprocess
import sys

def fetch_with_archive(url: str) -> str:
    proc = subprocess.run(
        ["./scripts/fetch_url.sh", url],
        capture_output=True,
        text=True,
    )
    if proc.returncode == 75:
        cap_url = next(
            line.split(":", 1)[1] for line in proc.stderr.splitlines()
            if line.startswith("ARCHIVE_CAPTCHA:")
        )
        raise RuntimeError(
            f"CAPTCHA required – please open {cap_url} in a browser and retry."
        )
    proc.check_returncode()
    return proc.stdout

if __name__ == "__main__":
    try:
        print(fetch_with_archive(sys.argv[1]))
    except RuntimeError as e:
        print(e, file=sys.stderr)
        sys.exit(1)

```

## Summary

- The **archive.today fallback** operates at Level 4 of the six-level retrieval cascade implemented in [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh).
- **CAPTCHA detection** relies on `_is_captcha_page` scanning HTML for keywords including "recaptcha" and "cloudflare.*challenge".
- Valid content must pass both challenge detection and content sanity verification via `_has_content` before text extraction proceeds.
- Upon detecting a CAPTCHA, the system emits `ARCHIVE_CAPTCHA:<url>` to stderr and exits with **code 75**, enabling upstream callers to initiate manual browser-based verification.
- This architecture prevents data corruption from incomplete fetches and provides clear, actionable feedback when automated retrieval encounters security barriers.

## Frequently Asked Questions

### What exit code indicates a CAPTCHA challenge from archive.today?

The script exits with **code 75** when archive.today returns a CAPTCHA page. This specific code is recognized by the Claude skill interface as indicating that human verification is required, allowing calling applications to distinguish between permanent fetch failures and temporary blocks solvable via manual intervention.

### How does the script distinguish between valid content and a CAPTCHA page?

The `_is_captcha_page` function in [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh) performs case-insensitive pattern matching against the HTTP response body, searching for strings such as "security check", "captcha", "recaptcha", and "cloudflare.*challenge". Simultaneously, `_has_content` verifies the response contains sufficient text volume to represent a genuine article rather than a placeholder or error page.

### Will the fallback return partial content if a CAPTCHA is detected?

No. When the script detects a CAPTCHA, it explicitly prevents content extraction and text conversion. It returns exit code 75 and diagnostic messages to stderr rather than risking the return of incomplete data, security check pages, or misleading placeholders that could contaminate the final notebook output.

### How should automation pipelines handle exit code 75?

Capture the specific archive URL from the stderr output where the script prints `ARCHIVE_CAPTCHA:<url>`, present this URL to the user for manual browser-based verification, and re-execute the fetch operation once the CAPTCHA is solved. Both the Bash and Python examples above demonstrate parsing the stderr stream to extract the verification URL and pause the pipeline appropriately.