How the Patent-Disclosure Crawler Determines Success or Failure in cnipa_crawler.py

The crawler determines success or failure through a deterministic seven-stage pipeline that validates browser initialization, HTTP response integrity, DOM structure, and pagination state, ultimately emitting standardized exit codes 0, 1, or 2 from the main execution block to signal full success, partial success, or critical failure.

The handsomestWei/patent-disclosure-skill repository automates patent searches through a Playwright-based crawling engine. At the heart of this system, skills/patent-search/tools/cnipa_crawler.py implements explicit success semantics that allow CI pipelines and downstream tooling to interpret crawl outcomes unambiguously.

The Seven-Stage Success/Failure Validation Pipeline

The script evaluates each crawl through a sequential validation process. Each stage contains specific failure points that trigger immediate termination or retry logic.

Stage 1: Browser Initialization Verification

The process begins by invoking launch_chromium() at line 33 of cnipa_crawler.py. This function establishes the Playwright browser context and navigates to the advanced search page. If the browser fails to launch—due to missing dependencies, memory constraints, or display issues—the resulting exception propagates upward uncaught, causing the interpreter to exit with code 1. This represents an infrastructure-level failure before any network requests occur.

Stage 2: HTTP Response Validation via JavaScript Bridge

Once the browser context exists, the crawler executes the _FETCH_RESULT_PAGE_JS snippet defined at line 15. This JavaScript code submits the search form and returns a JSON object containing ok, status, text, currentPage, and totalPages fields. The Python wrapper immediately inspects response["ok"].

If the HTTP status differs from 200 OK, or if the ok field is falsy, the crawler treats the attempt as a network-layer failure. According to the main execution block at line 424, these conditions precipitate a call to sys.exit(2), signaling a critical, non-recoverable error to the calling environment.

Stage 3: Result Page Structure Verification

After receiving an HTTP 200 response, the crawler must confirm the page actually contains search results. It executes _RESULT_PAGE_READY_JS at line 72, which inspects the document title and verifies the presence of result containers in the DOM.

If the page title indicates an error state, or if the expected result containers are absent (suggesting a server-side error page or unexpected layout change), the run marks itself as failed. This validation prevents the crawler from parsing empty or error pages as valid data.

Stage 4: Pagination Progress Monitoring

To detect silent failures or infinite loops, the script employs _PAGE_ADVANCED_JS at line 5. This routine captures the current page number and a fingerprint of the result set.

If subsequent iterations show identical page numbers or unchanged fingerprints, the crawler recognizes a stalled state. The logic enters a retry loop; if the stagnation persists beyond configured thresholds, the process transitions to the failure state defined in the exit handling logic.

Stage 5: Navigation Completion Logic

For each validated page, the crawler calls _FIND_NEXT_PAGE_JS at line 58 to locate the “next” navigation element in the DOM. Upon finding the element, it executes _CLICK_NEXT_PAGE_JS to advance the pagination.

The crawl determines success when it exhausts the configured page budget without finding a “next” link—indicating all available results were processed—or when it successfully processes the maximum requested pages. This graceful termination occurs only when all previous validation stages passed.

Stage 6: Hit Collection and Data Integrity Verification

As pages advance successfully, the crawler extracts patent records into EpubSearchHit objects. The parsing logic, primarily residing in cnipa_parse.py, raises exceptions for malformed HTML or missing required fields.

If every page through the navigation sequence returns valid, parseable data, the run_crawler function assembles a result dictionary with ok: True and populates the hits list. Any parsing exception propagates to the main loop, triggering the critical failure path.

Stage 7: Exit Code Classification and Process Termination

The definitive success or failure determination occurs in the if __name__ == "__main__": block at line 424. This wrapper implements a try/except structure that maps internal states to standard Unix exit codes:

  • sys.exit(0) – All pages processed without error; the result["ok"] flag is true.
  • sys.exit(1) – Non-critical errors occurred (e.g., missing optional metadata in search_config.py or partial data in some pages). The crawl delivered incomplete but usable results.
  • sys.exit(2) – Critical failure modes including HTTP 400+ responses, parsing exceptions from cnipa_parse.py, unrecoverable browser crashes from browser.py, or validation failures in Stages 2–4.

These exit codes provide the primary contract for external orchestration tools to determine success or failure without parsing log files.

Cross-Module Failure Propagation

While cnipa_crawler.py contains the arbitration logic, several companion modules influence the final determination:

  • cnipa_parse.py – Raises ValueError or parsing exceptions when HTML structure deviates from expected schemas. These bubble up to the main crawler loop and convert to exit 2.
  • search_config.py – Loads YAML search parameters. Invalid configurations trigger early termination with exit 1 before browser launch.
  • browser.py – Provides launch_chromium(). Failure to initialize the Playwright context propagates as an unhandled exception, typically resulting in exit 1.

This architecture ensures that failures in dependencies are captured and classified according to the crawler’s exit code semantics.

Usage Examples: Detecting Outcomes Programmatically

When importing the crawler as a library, inspect the returned dictionary:

from cnipa_crawler import run_crawler

config = load_search_config("my_config.yaml")
result = run_crawler(config)

if result["ok"]:
    print("✅ Crawl succeeded, got", len(result["hits"]), "hits.")
else:
    print("❌ Crawl failed – see logs for details.")

For command-line automation, check the shell exit status:


# Direct CLI invocation (exit code indicates outcome)

python -m skills.patent-search.tools.cnipa_crawler  # → 0 on success

echo $?  # 0 = success, 1 = partial, 2 = fatal error

The $? variable captures the exit code defined at line 424, allowing shell scripts to branch based on crawl success or failure deterministically.

Summary

  • Seven-stage validation: The crawler validates browser launch, HTTP responses, DOM readiness, pagination advancement, navigation completion, and data parsing before finalizing state.
  • Explicit exit codes: Line 424 of cnipa_crawler.py emits 0 (success), 1 (partial), or 2 (critical failure) to provide unambiguous machine-readable outcomes.
  • JavaScript integration: Constants like _FETCH_RESULT_PAGE_JS (line 15) and _PAGE_ADVANCED_JS (line 5) bridge Python and browser contexts to validate page state.
  • Cross-module errors: Exceptions from cnipa_parse.py, search_config.py, and browser.py propagate to the main loop and influence the final exit classification.
  • Reliable automation: The deterministic contract supports both programmatic Python imports and CLI invocations in CI/CD pipelines.

Frequently Asked Questions

How does the crawler distinguish between a partial success and a complete failure?

The main execution block at line 424 implements granular exit code logic. When the run_crawler function returns despite non-critical issues—such as missing optional configuration fields in search_config.py or isolated parsing errors in cnipa_parse.py—the script calls sys.exit(1). This indicates partial success: some data was retrieved but the dataset may be incomplete. Conversely, sys.exit(2) indicates critical failures like HTTP 400 errors, browser crashes from browser.py, or validation failures in the _RESULT_PAGE_READY_JS check, where the output is unreliable or empty.

What specific JavaScript checks determine if a result page is ready?

The crawler executes _RESULT_PAGE_READY_JS at line 72 to inspect the document state before parsing. This snippet validates the page title against expected patterns and confirms the presence of result container elements in the DOM. If the title suggests an error page (such as “Service Unavailable”) or if the container selectors return empty NodeLists, the script marks the run as failed before attempting patent extraction.

Why does the crawler use exit codes instead of return values for CLI usage?

Exit codes provide a universal interface for Unix-like environments and CI/CD systems that execute python -m skills.patent-search.tools.cnipa_crawler as a subprocess. While the Python API returns a dictionary with an ok boolean for programmatic use, the command-line interface must communicate status to shell scripts and automation tools. The standardized codes—0 for success, 1 for partial success, and 2 for failure—allow pipelines to use standard operators like && and || or inspect $? immediately after execution without parsing JSON or log files.

Where in the source code does the crawler detect stalled pagination?

The pagination watchdog logic resides in _PAGE_ADVANCED_JS at line 5. This JavaScript function captures the current page index and a content fingerprint. The Python loop compares these values across iterations; if the page number fails to increment or the fingerprint remains static despite a “next” click via _CLICK_NEXT_PAGE_JS, the crawler detects a stall. After exhausting configured retry attempts, it transitions to the failure state, ultimately resulting in sys.exit(2) from the main block.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →