How the patent-disclosure-skill Performs Prior Art Search Using CNIPA epub.cnipa.gov.cn

The patent-disclosure-skill automates prior art searches against CNIPA's epub.cnipa.gov.cn database by orchestrating a headless Chromium browser to fill the advanced search form, paginate through results using UI interactions and fallback HTTP requests, and generate structured Markdown reports with JSON payloads.

The patent-disclosure-skill repository (handsomestWei/patent-disclosure-skill) provides an open-source implementation for conducting automated prior art searches on the China National Intellectual Property Administration (CNIPA) publication database. This skill leverages Playwright to control a headless browser, enabling programmatic interaction with the advanced search interface at epub.cnipa.gov.cn/Advanced. Below is a detailed breakdown of how the system executes a complete prior art search using CNIPA epub.cnipa.gov.cn, from CLI argument parsing to final report generation.

Architecture Overview

The prior art search functionality is organized across three core modules: the CLI front-end, the browser automation engine, and the HTML parser. The primary components reside in skills/patent-search/tools/ and include cnipa_search.py for orchestration, cnipa_crawler.py for browser control, and cnipa_parse.py for data extraction.

Step-by-Step Execution Flow

CLI Entry and Configuration Loading

The process begins in skills/patent-search/tools/cnipa_search.py. The main function parses command-line arguments—including inventor names, applicants, titles, and IPC classes—and constructs a query dictionary. It loads user-specific settings from search_config.py to determine the maximum number of result pages to fetch via resolve_max_pages.

Browser Initialization

Once arguments are validated, the system initializes a headless Chromium instance. The launch_chromium function in skills/patent-search/tools/browser.py configures the browser with a custom User-Agent to mimic real traffic. In skills/patent-search/tools/cnipa_crawler.py, the _launch_browser method creates the browser context and prepares the automation session.

Form Submission on CNIPA Advanced Search Page

The crawler navigates to http://epub.cnipa.gov.cn/Advanced and waits for the search form to load using wait_for_epub_advanced_ready. The submit_advanced_query function then programmatically fills the form fields—inventor, applicant, title, and patent type checkboxes—before clicking the "查询" (Search) button. This triggers the results page load, which the crawler monitors for completion.

Result Pagination and Duplicate Detection

To retrieve comprehensive results, collect_result_pages in cnipa_crawler.py implements a robust pagination strategy. It first parses the initial page to determine total pages and page size via _read_pager_plan. For subsequent pages, it attempts to click the Next button using advance_to_next_result_page. If the UI blocks navigation, the system falls back to _fetch_next_fragment, which sends direct HTTP POST requests to retrieve result fragments. To prevent infinite loops, the crawler computes SHA-256 fingerprints of each page and detects duplicates before processing.

HTML Parsing and Data Extraction

Each retrieved page is passed to parse_search_result_html in skills/patent-search/tools/cnipa_parse.py. This function extracts structured data into EpubSearchHit objects, capturing publication numbers, titles, inventors, applicants, and document links. The crawler then merges hits using _merge_hits, deduplicating records based on publication number, application number, link, or title to ensure data integrity.

Filtering and Report Generation

Back in cnipa_search.py, the raw hits undergo filtering via filter_hits, which applies user-specified criteria for inventor names and applicant aliases while grouping multiple announcements for the same application. Finally, write_search_report in skills/patent-search/tools/emit_search_report.py generates a Markdown report saved to outputs/patent-search/ and prints a JSON payload to stdout containing metadata, query parameters, pagination statistics, and the filtered result set.

Command-Line Usage and Programmatic Integration

Users can execute searches directly from the command line or import the functionality into other Python modules.

To run a search for inventor "张三" limiting to 2 pages:

python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三" \
    --max-pages 2

The script outputs three formatted lines indicating the Markdown report path, JSON payload, and completion status:


EPUB_SEARCH_MD: outputs/patent-search/2026-09-06-...md
EPUB_SEARCH_JSON: {...JSON payload...}
EPUB_SEARCH_NOTE: pages=2 candidates=45 matched=30 complete=False …

For programmatic use within another skill:

from skills.patent-search.tools.cnipa_search import main as cnipa_main

# Execute search for specific inventor and applicant

exit_code = cnipa_main([
    "--inventor", "张三",
    "--applicant", "华为技术有限公司",
    "--complete"
])
print("Search finished with code:", exit_code)

Summary

  • The patent-disclosure-skill performs prior art searches by automating the CNIPA advanced search interface at epub.cnipa.gov.cn using Playwright and headless Chromium.
  • The CLI entry point in cnipa_search.py handles argument parsing, configuration loading, and result filtering.
  • Browser automation in cnipa_crawler.py manages form submission, pagination through both UI clicks and fallback HTTP POSTs, and SHA-256 duplicate detection.
  • HTML parsing in cnipa_parse.py converts raw result pages into structured EpubSearchHit objects.
  • The system returns exit code 0 for successful completion, 3 for partial crawls, and other non-zero codes for configuration or argument errors.
  • Output includes both human-readable Markdown reports and machine-readable JSON payloads suitable for downstream processing.

Frequently Asked Questions

What is the primary entry point for running a CNIPA prior art search?

The main entry point is skills/patent-search/tools/cnipa_search.py, which provides the main function. This module parses command-line arguments, loads configurations from search_config.py, initializes the crawler, and coordinates the entire search workflow before emitting the final report.

How does the skill handle pagination when the UI blocks navigation?

When the standard Next button becomes unresponsive, the crawler falls back to _fetch_next_fragment in cnipa_crawler.py, which sends direct HTTP POST requests to retrieve subsequent result pages. The system tracks page uniqueness using SHA-256 hashes to prevent infinite loops during this process.

What exit codes does the CLI return?

The CLI returns 0 when the crawl completes successfully and exhausts all requested pages. It returns 3 if the crawl stops before fetching all pages (incomplete). Other non-zero exit codes indicate argument validation failures or configuration errors detected during initialization.

Can the search be used programmatically from other Python modules?

Yes. The main function in cnipa_search.py is importable and accepts a list of string arguments. This allows other skills to execute searches by calling cnipa_main(["--inventor", "Name", "--complete"]) and processing the returned exit code to determine success or failure states.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →