How to Perform CNIPA Bibliographic Searches Using the Patent‑Search Skill
The CNIPA bibliographic search is implemented as a modular command‑line tool in the patent‑search skill that uses Playwright to crawl the CNIPA EPUB database, parses results into structured data, filters by inventor or applicant, and emits a Markdown report.
The patent‑search skill, part of the handsomestWei/patent-disclosure-skill repository, provides a dedicated toolchain for automating CNIPA bibliographic searches against the China National Intellectual Property Administration’s EPUB database. It combines headless browser automation with configurable filtering to produce structured patent reports without manual web navigation.
Architecture Overview
The CNIPA search toolchain consists of five decoupled components designed for testability and clear separation of concerns:
- CLI driver (
cnipa_search.py): Parses user arguments, loads configuration, and orchestrates the workflow. - Crawler (
cnipa_crawler.py): Manages a headless Chromium instance via Playwright to submit queries and paginate results. - Parser (
cnipa_parse.py): Extracts structured fields—such as title, publication number, and inventor list—from the CNIPA EPUB HTML. - Filter & aggregator (
filter_hitsincnipa_search.py): Applies optional inventor and applicant filters and merges duplicate publication records under the same application number. - Report writer (
emit_search_report.py): Converts aggregated hits into a Markdown report saved underoutputs/patent-search/.
The crawler can be unit‑tested in isolation using skills/patent-search/tests/test_cnipa_search.py, and the parser remains reusable for other front‑ends.
Step‑by‑Step Workflow
1. CLI Argument Handling
The entry point cnipa_search.py instantiates an argparse.ArgumentParser (lines 48‑77) that accepts inventor, applicant, title, IPC/LOC class, application number, publication number, patent type, and pagination controls (--max-pages and --complete).
2. Configuration Loading
The function search_config.load_search_config reads config.yaml (default max_pages: 3) and merges command‑line overrides via resolve_max_pages. This determines the page budget for the crawler.
3. Query Construction
The helper _query_fields (lines 82‑95) normalizes user inputs into a dictionary passed to the crawler.
4. Web Crawling with Playwright
cnipa_crawler.search_advanced launches Chromium (launch_chromium), fills the advanced query form, and paginates through results. JavaScript snippets (_RESULT_PAGE_READY_JS, _FETCH_RESULT_PAGE_JS, _FIND_NEXT_PAGE_JS) ensure the DOM is fully loaded before extraction and detect the final page.
5. HTML Parsing
For each fetched page, cnipa_parse.parse_search_result_html transforms HTML into a list of EpubSearchHit objects containing pub_number, application_number, inventors, applicant, link, and raw_html.
6. Filtering and Aggregation
The filter_hits function (lines 55‑115) optionally filters by inventor identity (normalize_identity) and applicant aliases (_matching_applicant). It also consolidates multiple publication records sharing the same application number.
7. Completeness Tracking
The helper completeness_note (lines 21‑45) generates a human‑readable status indicating whether the search reached the final page, how many pages were scanned, and whether a complete crawl was requested.
8. Report Generation
write_search_report in emit_search_report.py renders the aggregated rows as a Markdown table with hyperlinks and writes the file to outputs/patent-search/.
Usage Examples
Command‑Line Searches
Run a simple inventor‑only search using default pagination:
python skills/patent-search/tools/cnipa_search.py \
--inventor "张三"
Search by title and IPC class, limiting to two pages:
python skills/patent-search/tools/cnipa_search.py \
--title "数据处理" \
--class B01J20 \
--max-pages 2
Perform a complete crawl of all pages for a specific applicant:
python skills/patent-search/tools/cnipa_search.py \
--applicant "华为技术有限公司" \
--complete
Programmatic Usage
You can import the underlying functions for integration into larger pipelines:
from skills.patent-search.tools.cnipa_search import (
_build_parser, filter_hits, completeness_note
)
# Build an argparse.Namespace manually
args = _build_parser().parse_args([
"--inventor", "李四",
"--applicant", "中科院",
"--max-pages", "1"
])
# After crawling and parsing you obtain `hits` (list of EpubSearchHit)
filtered = filter_hits(hits, inventor=args.inventor, applicants=args.applicant)
note = completeness_note(
complete=False,
total_pages=5,
pages_scanned=1,
page_budget=args.max_pages,
pages_remaining=4,
has_next=True
)
print(filtered, note)
Summary
- The patent‑search skill provides a modular command‑line interface for CNIPA bibliographic searches located in
skills/patent-search/tools/. - Playwright drives the crawler (
cnipa_crawler.py) to navigate the CNIPA EPUB site, whilecnipa_parse.pyextracts structuredEpubSearchHitobjects. - The
filter_hitsfunction deduplicates records and filters by inventor or applicant identity before report generation. - Results are written as Markdown files to
outputs/patent-search/viaemit_search_report.py. - The tool supports both interactive CLI usage and programmatic import of its Python functions.
Frequently Asked Questions
What is the default pagination limit for CNIPA bibliographic searches?
The default limit is three pages, defined in config.yaml and loaded by search_config.load_search_config. You can override this with the --max-pages flag or request an unlimited crawl using the --complete flag.
How does the patent‑search skill handle duplicate patent records?
The filter_hits function in cnipa_search.py groups multiple publication records that share the same application number into a single consolidated entry, preventing redundant entries in the final report.
Can I use the CNIPA search module programmatically without the CLI?
Yes. The functions _build_parser, filter_hits, and completeness_note are importable from cnipa_search.py, allowing you to construct argument namespaces manually and process EpubSearchHit lists within Python scripts.
What dependencies are required to run the crawler?
The crawler requires Playwright and a headless Chromium browser instance. The skill uses JavaScript injection (_RESULT_PAGE_READY_JS, _FETCH_RESULT_PAGE_JS) to ensure the DOM is fully loaded before parsing each result page.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →