# How the Patent-Disclosure Crawler Determines Success or Failure in cnipa_crawler.py

> Discover how the patent-disclosure-skill crawler uses a seven-stage pipeline and exit codes 0 1 or 2 to determine success failure or partial success in cnipa_crawler.py.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: internals
- Published: 2026-09-04

---

**The crawler determines success or failure through a deterministic seven-stage pipeline that validates browser initialization, HTTP response integrity, DOM structure, and pagination state, ultimately emitting standardized exit codes 0, 1, or 2 from the main execution block to signal full success, partial success, or critical failure.**

The `handsomestWei/patent-disclosure-skill` repository automates patent searches through a Playwright-based crawling engine. At the heart of this system, [`skills/patent-search/tools/cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_crawler.py) implements explicit success semantics that allow CI pipelines and downstream tooling to interpret crawl outcomes unambiguously.

## The Seven-Stage Success/Failure Validation Pipeline

The script evaluates each crawl through a sequential validation process. Each stage contains specific failure points that trigger immediate termination or retry logic.

### Stage 1: Browser Initialization Verification

The process begins by invoking `launch_chromium()` at line 33 of [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py). This function establishes the Playwright browser context and navigates to the advanced search page. If the browser fails to launch—due to missing dependencies, memory constraints, or display issues—the resulting exception propagates upward uncaught, causing the interpreter to exit with code `1`. This represents an infrastructure-level failure before any network requests occur.

### Stage 2: HTTP Response Validation via JavaScript Bridge

Once the browser context exists, the crawler executes the **`_FETCH_RESULT_PAGE_JS`** snippet defined at line 15. This JavaScript code submits the search form and returns a JSON object containing `ok`, `status`, `text`, `currentPage`, and `totalPages` fields. The Python wrapper immediately inspects `response["ok"]`. 

If the HTTP status differs from `200 OK`, or if the `ok` field is falsy, the crawler treats the attempt as a network-layer failure. According to the main execution block at line 424, these conditions precipitate a call to **`sys.exit(2)`**, signaling a critical, non-recoverable error to the calling environment.

### Stage 3: Result Page Structure Verification

After receiving an HTTP 200 response, the crawler must confirm the page actually contains search results. It executes **`_RESULT_PAGE_READY_JS`** at line 72, which inspects the document title and verifies the presence of result containers in the DOM. 

If the page title indicates an error state, or if the expected result containers are absent (suggesting a server-side error page or unexpected layout change), the run marks itself as failed. This validation prevents the crawler from parsing empty or error pages as valid data.

### Stage 4: Pagination Progress Monitoring

To detect silent failures or infinite loops, the script employs **`_PAGE_ADVANCED_JS`** at line 5. This routine captures the current page number and a fingerprint of the result set. 

If subsequent iterations show identical page numbers or unchanged fingerprints, the crawler recognizes a stalled state. The logic enters a retry loop; if the stagnation persists beyond configured thresholds, the process transitions to the failure state defined in the exit handling logic.

### Stage 5: Navigation Completion Logic

For each validated page, the crawler calls **`_FIND_NEXT_PAGE_JS`** at line 58 to locate the “next” navigation element in the DOM. Upon finding the element, it executes **`_CLICK_NEXT_PAGE_JS`** to advance the pagination. 

The crawl determines **success** when it exhausts the configured page budget without finding a “next” link—indicating all available results were processed—or when it successfully processes the maximum requested pages. This graceful termination occurs only when all previous validation stages passed.

### Stage 6: Hit Collection and Data Integrity Verification

As pages advance successfully, the crawler extracts patent records into **`EpubSearchHit`** objects. The parsing logic, primarily residing in [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py), raises exceptions for malformed HTML or missing required fields. 

If every page through the navigation sequence returns valid, parseable data, the `run_crawler` function assembles a **`result`** dictionary with `ok: True` and populates the `hits` list. Any parsing exception propagates to the main loop, triggering the critical failure path.

### Stage 7: Exit Code Classification and Process Termination

The definitive success or failure determination occurs in the **`if __name__ == "__main__":`** block at line 424. This wrapper implements a `try/except` structure that maps internal states to standard Unix exit codes:

- **`sys.exit(0)`** – All pages processed without error; the `result["ok"]` flag is true.
- **`sys.exit(1)`** – Non-critical errors occurred (e.g., missing optional metadata in [`search_config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/search_config.py) or partial data in some pages). The crawl delivered incomplete but usable results.
- **`sys.exit(2)`** – Critical failure modes including HTTP 400+ responses, parsing exceptions from [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py), unrecoverable browser crashes from [`browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/browser.py), or validation failures in Stages 2–4.

These exit codes provide the primary contract for external orchestration tools to determine success or failure without parsing log files.

## Cross-Module Failure Propagation

While [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py) contains the arbitration logic, several companion modules influence the final determination:

- **[`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py)** – Raises `ValueError` or parsing exceptions when HTML structure deviates from expected schemas. These bubble up to the main crawler loop and convert to exit 2.
- **[`search_config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/search_config.py)** – Loads YAML search parameters. Invalid configurations trigger early termination with exit 1 before browser launch.
- **[`browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/browser.py)** – Provides `launch_chromium()`. Failure to initialize the Playwright context propagates as an unhandled exception, typically resulting in exit 1.

This architecture ensures that failures in dependencies are captured and classified according to the crawler’s exit code semantics.

## Usage Examples: Detecting Outcomes Programmatically

When importing the crawler as a library, inspect the returned dictionary:

```python
from cnipa_crawler import run_crawler

config = load_search_config("my_config.yaml")
result = run_crawler(config)

if result["ok"]:
    print("✅ Crawl succeeded, got", len(result["hits"]), "hits.")
else:
    print("❌ Crawl failed – see logs for details.")

```

For command-line automation, check the shell exit status:

```bash

# Direct CLI invocation (exit code indicates outcome)

python -m skills.patent-search.tools.cnipa_crawler  # → 0 on success

echo $?  # 0 = success, 1 = partial, 2 = fatal error

```

The `$?` variable captures the exit code defined at line 424, allowing shell scripts to branch based on crawl success or failure deterministically.

## Summary

- **Seven-stage validation**: The crawler validates browser launch, HTTP responses, DOM readiness, pagination advancement, navigation completion, and data parsing before finalizing state.
- **Explicit exit codes**: Line 424 of [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py) emits `0` (success), `1` (partial), or `2` (critical failure) to provide unambiguous machine-readable outcomes.
- **JavaScript integration**: Constants like `_FETCH_RESULT_PAGE_JS` (line 15) and `_PAGE_ADVANCED_JS` (line 5) bridge Python and browser contexts to validate page state.
- **Cross-module errors**: Exceptions from [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py), [`search_config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/search_config.py), and [`browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/browser.py) propagate to the main loop and influence the final exit classification.
- **Reliable automation**: The deterministic contract supports both programmatic Python imports and CLI invocations in CI/CD pipelines.

## Frequently Asked Questions

### How does the crawler distinguish between a partial success and a complete failure?

The main execution block at line 424 implements granular exit code logic. When the `run_crawler` function returns despite non-critical issues—such as missing optional configuration fields in [`search_config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/search_config.py) or isolated parsing errors in [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py)—the script calls **`sys.exit(1)`**. This indicates partial success: some data was retrieved but the dataset may be incomplete. Conversely, **`sys.exit(2)`** indicates critical failures like HTTP 400 errors, browser crashes from [`browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/browser.py), or validation failures in the `_RESULT_PAGE_READY_JS` check, where the output is unreliable or empty.

### What specific JavaScript checks determine if a result page is ready?

The crawler executes **`_RESULT_PAGE_READY_JS`** at line 72 to inspect the document state before parsing. This snippet validates the page title against expected patterns and confirms the presence of result container elements in the DOM. If the title suggests an error page (such as “Service Unavailable”) or if the container selectors return empty NodeLists, the script marks the run as failed before attempting patent extraction.

### Why does the crawler use exit codes instead of return values for CLI usage?

Exit codes provide a universal interface for Unix-like environments and CI/CD systems that execute `python -m skills.patent-search.tools.cnipa_crawler` as a subprocess. While the Python API returns a dictionary with an `ok` boolean for programmatic use, the command-line interface must communicate status to shell scripts and automation tools. The standardized codes—`0` for success, `1` for partial success, and `2` for failure—allow pipelines to use standard operators like `&&` and `||` or inspect `$?` immediately after execution without parsing JSON or log files.

### Where in the source code does the crawler detect stalled pagination?

The pagination watchdog logic resides in **`_PAGE_ADVANCED_JS`** at line 5. This JavaScript function captures the current page index and a content fingerprint. The Python loop compares these values across iterations; if the page number fails to increment or the fingerprint remains static despite a “next” click via **`_CLICK_NEXT_PAGE_JS`**, the crawler detects a stall. After exhausting configured retry attempts, it transitions to the failure state, ultimately resulting in **`sys.exit(2)`** from the main block.