How to Perform CNIPA Bibliographic Searches Using the Patent‑Search Skill

The CNIPA bibliographic search is implemented as a modular command‑line tool in the patent‑search skill that uses Playwright to crawl the CNIPA EPUB database, parses results into structured data, filters by inventor or applicant, and emits a Markdown report.

The patent‑search skill, part of the handsomestWei/patent-disclosure-skill repository, provides a dedicated toolchain for automating CNIPA bibliographic searches against the China National Intellectual Property Administration’s EPUB database. It combines headless browser automation with configurable filtering to produce structured patent reports without manual web navigation.

Architecture Overview

The CNIPA search toolchain consists of five decoupled components designed for testability and clear separation of concerns:

  • CLI driver (cnipa_search.py): Parses user arguments, loads configuration, and orchestrates the workflow.
  • Crawler (cnipa_crawler.py): Manages a headless Chromium instance via Playwright to submit queries and paginate results.
  • Parser (cnipa_parse.py): Extracts structured fields—such as title, publication number, and inventor list—from the CNIPA EPUB HTML.
  • Filter & aggregator (filter_hits in cnipa_search.py): Applies optional inventor and applicant filters and merges duplicate publication records under the same application number.
  • Report writer (emit_search_report.py): Converts aggregated hits into a Markdown report saved under outputs/patent-search/.

The crawler can be unit‑tested in isolation using skills/patent-search/tests/test_cnipa_search.py, and the parser remains reusable for other front‑ends.

Step‑by‑Step Workflow

1. CLI Argument Handling

The entry point cnipa_search.py instantiates an argparse.ArgumentParser (lines 48‑77) that accepts inventor, applicant, title, IPC/LOC class, application number, publication number, patent type, and pagination controls (--max-pages and --complete).

2. Configuration Loading

The function search_config.load_search_config reads config.yaml (default max_pages: 3) and merges command‑line overrides via resolve_max_pages. This determines the page budget for the crawler.

3. Query Construction

The helper _query_fields (lines 82‑95) normalizes user inputs into a dictionary passed to the crawler.

4. Web Crawling with Playwright

cnipa_crawler.search_advanced launches Chromium (launch_chromium), fills the advanced query form, and paginates through results. JavaScript snippets (_RESULT_PAGE_READY_JS, _FETCH_RESULT_PAGE_JS, _FIND_NEXT_PAGE_JS) ensure the DOM is fully loaded before extraction and detect the final page.

5. HTML Parsing

For each fetched page, cnipa_parse.parse_search_result_html transforms HTML into a list of EpubSearchHit objects containing pub_number, application_number, inventors, applicant, link, and raw_html.

6. Filtering and Aggregation

The filter_hits function (lines 55‑115) optionally filters by inventor identity (normalize_identity) and applicant aliases (_matching_applicant). It also consolidates multiple publication records sharing the same application number.

7. Completeness Tracking

The helper completeness_note (lines 21‑45) generates a human‑readable status indicating whether the search reached the final page, how many pages were scanned, and whether a complete crawl was requested.

8. Report Generation

write_search_report in emit_search_report.py renders the aggregated rows as a Markdown table with hyperlinks and writes the file to outputs/patent-search/.

Usage Examples

Command‑Line Searches

Run a simple inventor‑only search using default pagination:

python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三"

Search by title and IPC class, limiting to two pages:

python skills/patent-search/tools/cnipa_search.py \
    --title "数据处理" \
    --class B01J20 \
    --max-pages 2

Perform a complete crawl of all pages for a specific applicant:

python skills/patent-search/tools/cnipa_search.py \
    --applicant "华为技术有限公司" \
    --complete

Programmatic Usage

You can import the underlying functions for integration into larger pipelines:

from skills.patent-search.tools.cnipa_search import (
    _build_parser, filter_hits, completeness_note
)

# Build an argparse.Namespace manually

args = _build_parser().parse_args([
    "--inventor", "李四",
    "--applicant", "中科院",
    "--max-pages", "1"
])

# After crawling and parsing you obtain `hits` (list of EpubSearchHit)

filtered = filter_hits(hits, inventor=args.inventor, applicants=args.applicant)
note = completeness_note(
    complete=False,
    total_pages=5,
    pages_scanned=1,
    page_budget=args.max_pages,
    pages_remaining=4,
    has_next=True
)
print(filtered, note)

Summary

  • The patent‑search skill provides a modular command‑line interface for CNIPA bibliographic searches located in skills/patent-search/tools/.
  • Playwright drives the crawler (cnipa_crawler.py) to navigate the CNIPA EPUB site, while cnipa_parse.py extracts structured EpubSearchHit objects.
  • The filter_hits function deduplicates records and filters by inventor or applicant identity before report generation.
  • Results are written as Markdown files to outputs/patent-search/ via emit_search_report.py.
  • The tool supports both interactive CLI usage and programmatic import of its Python functions.

Frequently Asked Questions

What is the default pagination limit for CNIPA bibliographic searches?

The default limit is three pages, defined in config.yaml and loaded by search_config.load_search_config. You can override this with the --max-pages flag or request an unlimited crawl using the --complete flag.

How does the patent‑search skill handle duplicate patent records?

The filter_hits function in cnipa_search.py groups multiple publication records that share the same application number into a single consolidated entry, preventing redundant entries in the final report.

Can I use the CNIPA search module programmatically without the CLI?

Yes. The functions _build_parser, filter_hits, and completeness_note are importable from cnipa_search.py, allowing you to construct argument namespaces manually and process EpubSearchHit lists within Python scripts.

What dependencies are required to run the crawler?

The crawler requires Playwright and a headless Chromium browser instance. The skill uses JavaScript injection (_RESULT_PAGE_READY_JS, _FETCH_RESULT_PAGE_JS) to ensure the DOM is fully loaded before parsing each result page.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →