How the Patent-Search Sub-Skill Performs CNIPA Patent Searches: A Complete Technical Guide
The patent-search sub-skill queries CNIPA (China National Intellectual Property Administration) through an 8-step Playwright-driven pipeline that scrapes the EPub Advanced Search portal, normalizes applicant identities, filters duplicate records, and outputs structured Markdown/JSON reports.
This article explains exactly how the handsomestWei/patent-disclosure-skill repository implements automated Chinese patent searches against the official CNIPA database. The implementation isolates browser automation from business logic, making the system both robust and testable.
Overview of the CNIPA Search Architecture
The patent-search sub-skill follows a data-flow pipeline: CLI → Configuration → Playwright Crawler → HTML Parser → Normalizer/Filter → Payload Assembly → Report Generation. Each stage is encapsulated in dedicated modules, allowing independent testing and maintenance.
Core Pipeline Components
| Stage | File | Primary Function |
|---|---|---|
| Argument parsing | cnipa_search.py |
_build_parser() at lines 48-79 |
| Identity normalization | cnipa_search.py |
normalize_identity and _matching_applicant at lines 40-52 |
| Configuration loading | search_config.py |
load_search_config and resolve_max_pages at lines 20-46 |
| Browser automation | cnipa_crawler.py |
search_advanced at lines 260-340 |
| HTML extraction | cnipa_parse.py |
parse_search_result_html at lines 349-419 |
| Result filtering | cnipa_search.py |
filter_hits at lines 55-115 |
| Report generation | emit_search_report.py |
write_search_report at lines 57-85 |
Step 1: CLI Argument Parsing and Query Building
The search begins in cnipa_search.py with _build_parser() constructing a comprehensive argument interface. Users can specify:
- Inventor names (
--inventor) - Applicant/assignee names (
--applicant) - Patent title keywords (
--title) - IPC/LOC classification codes (
--class) - Application or publication numbers (
--application-number,--publication-number) - Patent type filter (
--type:invention,utility_model,design,all) - Pagination control (
--max-pagesor--completefor exhaustive search)
python skills/patent-search/tools/cnipa_search.py \
--inventor "张三" \
--type invention \
--max-pages 2
For exhaustive searches across all result pages:
python skills/patent-search/tools/cnipa_search.py \
--applicant "华为技术有限公司" \
--class B01J20 \
--complete
Step 2: Configuration and Pagination Limits
The search_config.py module loads runtime parameters from config.yaml via load_search_config(). The resolve_max_pages() function reconciles:
- Default page limits from configuration
- Explicit
--max-pagesuser override - The
--completeflag (which sets an effectively unlimited budget)
This separation ensures consistent pagination behavior across all CNIPA search operations.
Step 3: Playwright-Driven Browser Automation
The heavy lifting occurs in cnipa_crawler.py through the search_advanced() method (lines 260-340). This component:
- Launches a headless Chromium instance via Playwright
- Navigates to
http://epub.cnipa.gov.cn/Advanced - Populates the Advanced Search form with query parameters
- Applies the patent type filter (invention, utility model, or design)
- Submits the query and waits for result rendering
- Iterates through pagination until reaching the page budget or final result page
This browser automation approach is necessary because CNIPA's EPub portal uses JavaScript-heavy interfaces that resist simple HTTP requests.
Step 4: HTML Parsing and Data Extraction
Raw HTML responses are processed by cnipa_parse.py::parse_search_result_html() (lines 349-419), which extracts each search hit into an EpubSearchHit dataclass containing:
- Inventor list
- Applicant/assignee name
- Publication number
- Application number
- IPC (International Patent Classification) codes
- LOC (Locarno Classification) codes for designs
- Direct link to the detailed record page
The parser handles CNIPA's specific HTML structure and encoding quirks, isolating DOM traversal from crawling logic.
Step 5: Identity Normalization and Hit Filtering
Before final output, results pass through cnipa_search.py::filter_hits() (lines 55-115), which implements two critical business rules:
normalize_identity(): Normalizes Unicode characters, collapses whitespace, and applies case-folding to handle inconsistent name representations in Chinese patent data.
_matching_applicant(): Performs substring matching against applicant aliases, accommodating subsidiaries, historical name variations, and partial matches.
The filter also de-duplicates records that share the same application number (common when publication and application records both appear) and decorates each hit with:
identity_status: Match confidence indicatormatched_applicant: Specifically matched alias from the query set
Step 6: Payload Assembly and Report Generation
The cnipa_search.py::main() function (lines 97-82) assembles a comprehensive JSON payload containing:
- Original query parameters
- Pagination metadata (pages scanned, completeness flag)
- Filtered and annotated hit list
- Processing notes
emit_search_report.py::write_search_report() (lines 57-85) then persists dual outputs:
- Markdown report for human review
- JSON copy for downstream skill consumption
Both files are written to outputs/patent-search/ with timestamped filenames.
Programmatic Integration
The sub-skill supports direct Python import for integration with larger automation workflows:
from skills.patent-search.tools.cnipa_search import main as cnipa_main
# Emulate CLI arguments
args = [
"--inventor", "李四",
"--type", "utility_model",
"--max-pages", "1"
]
exit_code = cnipa_main(args)
print(f"Search finished with exit code {exit_code}")
Inspecting generated results:
import json
from pathlib import Path
report_path = Path("outputs/patent-search") / "search_2023-09-05_14-30-12.json"
payload = json.loads(report_path.read_text(encoding="utf-8"))
print(f"Found {payload['matched_count']} matched records")
Key Implementation Files
| File Path | Responsibility |
|---|---|
skills/patent-search/tools/cnipa_search.py |
CLI entry point, orchestration, filtering |
skills/patent-search/tools/cnipa_crawler.py |
Playwright browser automation |
skills/patent-search/tools/cnipa_parse.py |
HTML parsing, EpubSearchHit extraction |
skills/patent-search/tools/search_config.py |
YAML configuration, pagination resolution |
skills/patent-search/tools/emit_search_report.py |
Markdown/JSON report formatting |
skills/patent-search/tools/patent_type.py |
Patent type string normalization |
skills/patent-search/tools/stdio_utf8.py |
UTF-8 I/O handling for cross-platform stability |
Design Strengths of the CNIPA Search Implementation
- Separation of concerns: Browser I/O isolated in
cnipa_crawler.py; pure Python logic in other modules enables unit testing without Playwright dependencies - Configurable pagination:
--max-pagesand--completeflags balance thoroughness against runtime - Identity-aware filtering: Unicode normalization and alias matching handle real-world data quality issues in Chinese patent records
- Dual output formats: Human-readable Markdown plus machine-parseable JSON serves both manual review and downstream automation
Summary
- The patent-search sub-skill queries CNIPA through an 8-stage pipeline ending in structured Markdown/JSON reports
- Playwright-driven crawling in
cnipa_crawler.pyhandles JavaScript-heavy EPub portal interactions - Identity normalization (
normalize_identity,_matching_applicant) manages Chinese name variations and applicant aliases - Configurable pagination via
--max-pagesor--completeflags controls search depth - Modular architecture separates untestable browser automation from pure Python business logic
Frequently Asked Questions
What is CNIPA and why does this sub-skill target it specifically?
CNIPA (China National Intellectual Property Administration) is the official Chinese patent office. The sub-skill targets CNIPA's EPub portal (epub.cnipa.gov.cn) because it provides the authoritative, publicly accessible database of Chinese invention patents, utility models, and design patents. The implementation specifically uses the Advanced Search interface to support complex multi-field queries.
Why does the implementation use Playwright instead of direct API calls?
CNIPA's EPub portal is a JavaScript-rendered web application without a documented public API. Direct HTTP requests return incomplete markup or anti-bot challenges. Playwright automates a real Chromium browser, executing JavaScript and rendering the full DOM—this is the only reliable method to extract structured patent data from CNIPA as of the implementation date.
How does the identity matching handle Chinese name variations?
The normalize_identity() function in cnipa_search.py applies Unicode normalization (NFKC), whitespace collapsing, and case-folding to handle encoding inconsistencies. The _matching_applicant() function then performs substring matching against configured aliases, accommodating corporate name variations like "华为技术有限公司" matching "华为" or "Huawei."
Can the search run without headless browser automation?
No. According to the source code in handsomestWei/patent-disclosure-skill, the CNIPA EPub portal requires JavaScript execution for form submission and result rendering. All patent data retrieval flows through cnipa_crawler.py::search_advanced(), which depends on Playwright. There is no REST API fallback in the current implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →