How Patent-Disclosure-Skill Manages Chinese Patent Lifecycle: A Technical Deep Dive
The patent-disclosure-skill automates the complete Chinese patent lifecycle by chaining specialized tools that query the CNIPA portal, crawl HTML results using Playwright, parse structured data, merge duplicate publication records, and generate consumable Markdown and JSON reports for downstream disclosure generation.
The patent-disclosure-skill repository implements an end-to-end pipeline for discovering and reporting Chinese patent data from the China National Intellectual Property Administration (CNIPA). By orchestrating browser automation, data normalization, and dual-format reporting, this open-source skill enables automated tracking of patent publications from raw web acquisition to final disclosure documents.
The Six-Stage CNIPA Pipeline Architecture
Managing the Chinese patent lifecycle requires handling dynamic web interfaces, inconsistent HTML structures, and duplicate publication records. The skill solves this through six coordinated stages, each handled by dedicated modules in skills/patent-search/tools/.
1. Search Query Construction with CLI Normalization
The workflow initiates in skills/patent-search/tools/cnipa_search.py, which exposes a command-line interface accepting:
--inventor– Individual inventor names (e.g., "张三")--applicant– Corporate assignees--title– Title keywords--ipc– International Patent Classification codes--type– Patent category (invention, utility model, or design)--max-pages– Pagination budget limits
This module validates patent type selections, normalizes Unicode inputs, and loads default paging configurations before triggering the crawler.
2. Playwright-Based Web Crawling
The skills/patent-search/tools/cnipa_crawler.py component manages browser automation using Playwright (dependency managed via requirements-cnipa.txt). Unlike static HTTP clients, this crawler:
- Opens the CNIPA "Advanced Search" SPA (Single Page Application)
- Fills dynamic form fields with search parameters
- Handles JavaScript-rendered pagination
- Respects the
--max-pagesbudget or uses the--completeflag to exhaustively retrieve all results when total page counts are known
The crawler emits raw HTML for each result page, capturing content invisible to traditional scraping tools.
3. HTML Parsing into Structured Objects
Raw HTML feeds into skills/patent-search/tools/cnipa_parse.py, which instantiates EpubSearchHit objects containing:
- Application numbers
- Publication numbers and dates
- Inventor and applicant arrays
- Document deep-links
The parser exposes application_number_for_epub_query for query string normalization and extracts the total result count from CNIPA's pagination metadata.
4. Application-Level Deduplication
Within cnipa_search.py, the filter_hits function performs critical data cleansing. It normalizes identity strings, applies optional inventor/applicant filters, and deduplicates records by application_number. Rather than listing each publication separately, the function aggregates all publication_records under their parent application entry, yielding a compact list where each invention appears exactly once regardless of multiple publication phases.
5. Dual-Format Report Emission
The skills/patent-search/tools/emit_search_report.py module writes timestamped outputs to outputs/patent-search/:
- Markdown – Human-readable tables with query context and patent summaries
- JSON – Machine-readable structures containing full metadata arrays
Both formats include query timestamps, paging statistics, completeness flags, and concise status lines to stderr (format: pages=5 candidates=32 matched=12 complete=true).
6. Disclosure Document Integration
The pipeline culminates in skills/patent-disclosure/tools/crawl/cnipa_epub_search.py, which imports the search utilities to feed reports into disclosure generators. This integration transforms the JSON hit lists into formal disclosure documents (Markdown → DOCX), completing the lifecycle from raw CNIPA data to legally formatted artifacts.
Practical Implementation Examples
Command-Line Search Execution
Invoke the standard workflow directly from the terminal:
python skills/patent-search/tools/cnipa_search.py \
--inventor "李四" \
--applicant "华为技术有限公司" \
--type invention \
--max-pages 5
This executes a filtered search for invention patents, retrieves up to five result pages, and writes both Markdown and JSON reports to the outputs directory.
Programmatic Python API
Embed the search pipeline in Python applications:
from skills.patent-search.tools.cnipa_search import main
# Search for patents by inventor "王五" with strict pagination
exit_code = main([
"--inventor", "王五",
"--max-pages", "2"
])
print(f"Search finished with exit code {exit_code}")
Processing Generated Reports
Consume the JSON output for custom analytics or visualization:
import json
from pathlib import Path
report_json = Path("outputs/patent-search/2024-05-01_14-23-10.json")
data = json.loads(report_json.read_text(encoding="utf-8"))
# Access deduplicated application records
for hit in data["hits"]:
pub_count = len(hit["publication_records"])
print(f"Application {hit['application_number']} – {pub_count} publications")
Key Source Files
| File | Primary Function |
|---|---|
skills/patent-search/tools/cnipa_search.py |
CLI orchestrator; query building, filter_hits deduplication |
skills/patent-search/tools/cnipa_crawler.py |
Playwright browser automation for CNIPA navigation |
skills/patent-search/tools/cnipa_parse.py |
HTML extraction and EpubSearchHit instantiation |
skills/patent-search/tools/emit_search_report.py |
Markdown/JSON report generation with metadata |
skills/patent-disclosure/tools/crawl/cnipa_epub_search.py |
Disclosure skill integration for document production |
Summary
- The skill manages the Chinese patent lifecycle through a six-stage pipeline: query construction, Playwright crawling, HTML parsing, deduplication, report emission, and disclosure integration.
- Playwright in
cnipa_crawler.pyhandles JavaScript-rendered CNIPA content unreachable by static requests. - The
filter_hitsfunction aggregates multiple publications under singleapplication_numberentries to eliminate redundancy. - Reports generate in dual formats (Markdown for humans, JSON for machines) with comprehensive metadata and pagination statistics.
- Modular architecture allows the
patent-disclosureskill to consume search outputs and produce formal DOCX disclosure documents.
Frequently Asked Questions
How does the skill handle CNIPA's JavaScript-heavy interface?
The implementation uses Playwright to drive a headless browser. The cnipa_crawler.py module renders the CNIPA Advanced Search page dynamically, fills interactive form fields, and extracts HTML after JavaScript execution completes. This approach succeeds where traditional requests or urllib would fail to retrieve the dynamic result tables.
What prevents duplicate patents from appearing in reports?
The filter_hits function normalizes identity strings and groups records by application number. Instead of listing each publication date separately, it creates a single entry per invention containing an aggregated publication_records array. This deduplication ensures that continuation publications or correction notices do not create redundant entries in the final output.
Can searches be restricted by patent type or result volume?
Yes. The CLI accepts a --type argument (invention, utility model, or design) to filter patent categories. Pagination is controlled via --max-pages for budgeted searches, while the --complete flag forces comprehensive retrieval when the total result count is known and full data capture is required.
Where are search reports stored and how are they structured?
The emit_search_report.py module writes timestamped files to outputs/patent-search/ in two formats: Markdown for human review and JSON for programmatic consumption. Both include query parameters, execution timestamps, completeness indicators, and deduplicated hit arrays, with status summaries printed to stderr during execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →