How Patent-Disclosure-Skill Manages Chinese Patent Lifecycle: A Technical Deep Dive

The patent-disclosure-skill automates the complete Chinese patent lifecycle by chaining specialized tools that query the CNIPA portal, crawl HTML results using Playwright, parse structured data, merge duplicate publication records, and generate consumable Markdown and JSON reports for downstream disclosure generation.

The patent-disclosure-skill repository implements an end-to-end pipeline for discovering and reporting Chinese patent data from the China National Intellectual Property Administration (CNIPA). By orchestrating browser automation, data normalization, and dual-format reporting, this open-source skill enables automated tracking of patent publications from raw web acquisition to final disclosure documents.

The Six-Stage CNIPA Pipeline Architecture

Managing the Chinese patent lifecycle requires handling dynamic web interfaces, inconsistent HTML structures, and duplicate publication records. The skill solves this through six coordinated stages, each handled by dedicated modules in skills/patent-search/tools/.

1. Search Query Construction with CLI Normalization

The workflow initiates in skills/patent-search/tools/cnipa_search.py, which exposes a command-line interface accepting:

  • --inventor – Individual inventor names (e.g., "张三")
  • --applicant – Corporate assignees
  • --title – Title keywords
  • --ipc – International Patent Classification codes
  • --type – Patent category (invention, utility model, or design)
  • --max-pages – Pagination budget limits

This module validates patent type selections, normalizes Unicode inputs, and loads default paging configurations before triggering the crawler.

2. Playwright-Based Web Crawling

The skills/patent-search/tools/cnipa_crawler.py component manages browser automation using Playwright (dependency managed via requirements-cnipa.txt). Unlike static HTTP clients, this crawler:

  • Opens the CNIPA "Advanced Search" SPA (Single Page Application)
  • Fills dynamic form fields with search parameters
  • Handles JavaScript-rendered pagination
  • Respects the --max-pages budget or uses the --complete flag to exhaustively retrieve all results when total page counts are known

The crawler emits raw HTML for each result page, capturing content invisible to traditional scraping tools.

3. HTML Parsing into Structured Objects

Raw HTML feeds into skills/patent-search/tools/cnipa_parse.py, which instantiates EpubSearchHit objects containing:

  • Application numbers
  • Publication numbers and dates
  • Inventor and applicant arrays
  • Document deep-links

The parser exposes application_number_for_epub_query for query string normalization and extracts the total result count from CNIPA's pagination metadata.

4. Application-Level Deduplication

Within cnipa_search.py, the filter_hits function performs critical data cleansing. It normalizes identity strings, applies optional inventor/applicant filters, and deduplicates records by application_number. Rather than listing each publication separately, the function aggregates all publication_records under their parent application entry, yielding a compact list where each invention appears exactly once regardless of multiple publication phases.

5. Dual-Format Report Emission

The skills/patent-search/tools/emit_search_report.py module writes timestamped outputs to outputs/patent-search/:

  • Markdown – Human-readable tables with query context and patent summaries
  • JSON – Machine-readable structures containing full metadata arrays

Both formats include query timestamps, paging statistics, completeness flags, and concise status lines to stderr (format: pages=5 candidates=32 matched=12 complete=true).

6. Disclosure Document Integration

The pipeline culminates in skills/patent-disclosure/tools/crawl/cnipa_epub_search.py, which imports the search utilities to feed reports into disclosure generators. This integration transforms the JSON hit lists into formal disclosure documents (Markdown → DOCX), completing the lifecycle from raw CNIPA data to legally formatted artifacts.

Practical Implementation Examples

Command-Line Search Execution

Invoke the standard workflow directly from the terminal:

python skills/patent-search/tools/cnipa_search.py \
    --inventor "李四" \
    --applicant "华为技术有限公司" \
    --type invention \
    --max-pages 5

This executes a filtered search for invention patents, retrieves up to five result pages, and writes both Markdown and JSON reports to the outputs directory.

Programmatic Python API

Embed the search pipeline in Python applications:

from skills.patent-search.tools.cnipa_search import main

# Search for patents by inventor "王五" with strict pagination

exit_code = main([
    "--inventor", "王五",
    "--max-pages", "2"
])
print(f"Search finished with exit code {exit_code}")

Processing Generated Reports

Consume the JSON output for custom analytics or visualization:

import json
from pathlib import Path

report_json = Path("outputs/patent-search/2024-05-01_14-23-10.json")
data = json.loads(report_json.read_text(encoding="utf-8"))

# Access deduplicated application records

for hit in data["hits"]:
    pub_count = len(hit["publication_records"])
    print(f"Application {hit['application_number']} – {pub_count} publications")

Key Source Files

File Primary Function
skills/patent-search/tools/cnipa_search.py CLI orchestrator; query building, filter_hits deduplication
skills/patent-search/tools/cnipa_crawler.py Playwright browser automation for CNIPA navigation
skills/patent-search/tools/cnipa_parse.py HTML extraction and EpubSearchHit instantiation
skills/patent-search/tools/emit_search_report.py Markdown/JSON report generation with metadata
skills/patent-disclosure/tools/crawl/cnipa_epub_search.py Disclosure skill integration for document production

Summary

  • The skill manages the Chinese patent lifecycle through a six-stage pipeline: query construction, Playwright crawling, HTML parsing, deduplication, report emission, and disclosure integration.
  • Playwright in cnipa_crawler.py handles JavaScript-rendered CNIPA content unreachable by static requests.
  • The filter_hits function aggregates multiple publications under single application_number entries to eliminate redundancy.
  • Reports generate in dual formats (Markdown for humans, JSON for machines) with comprehensive metadata and pagination statistics.
  • Modular architecture allows the patent-disclosure skill to consume search outputs and produce formal DOCX disclosure documents.

Frequently Asked Questions

How does the skill handle CNIPA's JavaScript-heavy interface?

The implementation uses Playwright to drive a headless browser. The cnipa_crawler.py module renders the CNIPA Advanced Search page dynamically, fills interactive form fields, and extracts HTML after JavaScript execution completes. This approach succeeds where traditional requests or urllib would fail to retrieve the dynamic result tables.

What prevents duplicate patents from appearing in reports?

The filter_hits function normalizes identity strings and groups records by application number. Instead of listing each publication date separately, it creates a single entry per invention containing an aggregated publication_records array. This deduplication ensures that continuation publications or correction notices do not create redundant entries in the final output.

Can searches be restricted by patent type or result volume?

Yes. The CLI accepts a --type argument (invention, utility model, or design) to filter patent categories. Pagination is controlled via --max-pages for budgeted searches, while the --complete flag forces comprehensive retrieval when the total result count is known and full data capture is required.

Where are search reports stored and how are they structured?

The emit_search_report.py module writes timestamped files to outputs/patent-search/ in two formats: Markdown for human review and JSON for programmatic consumption. Both include query parameters, execution timestamps, completeness indicators, and deduplicated hit arrays, with status summaries printed to stderr during execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →