# How Patent-Disclosure-Skill Manages Chinese Patent Lifecycle: A Technical Deep Dive

> Discover how patent-disclosure-skill automates the Chinese patent lifecycle. Explore its technical deep dive into CNIPA portal queries, Playwright crawling, data parsing, and report generation.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: deep-dive
- Published: 2026-09-05

---

**The patent-disclosure-skill automates the complete Chinese patent lifecycle by chaining specialized tools that query the CNIPA portal, crawl HTML results using Playwright, parse structured data, merge duplicate publication records, and generate consumable Markdown and JSON reports for downstream disclosure generation.**

The [patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill) repository implements an end-to-end pipeline for discovering and reporting Chinese patent data from the China National Intellectual Property Administration (CNIPA). By orchestrating browser automation, data normalization, and dual-format reporting, this open-source skill enables automated tracking of patent publications from raw web acquisition to final disclosure documents.

## The Six-Stage CNIPA Pipeline Architecture

Managing the **Chinese patent lifecycle** requires handling dynamic web interfaces, inconsistent HTML structures, and duplicate publication records. The skill solves this through six coordinated stages, each handled by dedicated modules in `skills/patent-search/tools/`.

### 1. Search Query Construction with CLI Normalization

The workflow initiates in [`skills/patent-search/tools/cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_search.py), which exposes a command-line interface accepting:

- `--inventor` – Individual inventor names (e.g., "张三")
- `--applicant` – Corporate assignees
- `--title` – Title keywords
- `--ipc` – International Patent Classification codes
- `--type` – Patent category (invention, utility model, or design)
- `--max-pages` – Pagination budget limits

This module validates patent type selections, normalizes Unicode inputs, and loads default paging configurations before triggering the crawler.

### 2. Playwright-Based Web Crawling

The [`skills/patent-search/tools/cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_crawler.py) component manages browser automation using **Playwright** (dependency managed via [`requirements-cnipa.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements-cnipa.txt)). Unlike static HTTP clients, this crawler:

- Opens the CNIPA "Advanced Search" SPA (Single Page Application)
- Fills dynamic form fields with search parameters
- Handles JavaScript-rendered pagination
- Respects the `--max-pages` budget or uses the `--complete` flag to exhaustively retrieve all results when total page counts are known

The crawler emits raw HTML for each result page, capturing content invisible to traditional scraping tools.

### 3. HTML Parsing into Structured Objects

Raw HTML feeds into [`skills/patent-search/tools/cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_parse.py), which instantiates **`EpubSearchHit`** objects containing:

- Application numbers
- Publication numbers and dates
- Inventor and applicant arrays
- Document deep-links

The parser exposes `application_number_for_epub_query` for query string normalization and extracts the total result count from CNIPA's pagination metadata.

### 4. Application-Level Deduplication

Within [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py), the **`filter_hits`** function performs critical data cleansing. It normalizes identity strings, applies optional inventor/applicant filters, and **deduplicates** records by `application_number`. Rather than listing each publication separately, the function aggregates all `publication_records` under their parent application entry, yielding a compact list where each invention appears exactly once regardless of multiple publication phases.

### 5. Dual-Format Report Emission

The [`skills/patent-search/tools/emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/emit_search_report.py) module writes timestamped outputs to `outputs/patent-search/`:

- **Markdown** – Human-readable tables with query context and patent summaries
- **JSON** – Machine-readable structures containing full metadata arrays

Both formats include query timestamps, paging statistics, completeness flags, and concise status lines to `stderr` (format: `pages=5 candidates=32 matched=12 complete=true`).

### 6. Disclosure Document Integration

The pipeline culminates in [`skills/patent-disclosure/tools/crawl/cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/crawl/cnipa_epub_search.py), which imports the search utilities to feed reports into disclosure generators. This integration transforms the JSON hit lists into formal disclosure documents (Markdown → DOCX), completing the lifecycle from raw CNIPA data to legally formatted artifacts.

## Practical Implementation Examples

### Command-Line Search Execution

Invoke the standard workflow directly from the terminal:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --inventor "李四" \
    --applicant "华为技术有限公司" \
    --type invention \
    --max-pages 5

```

This executes a filtered search for invention patents, retrieves up to five result pages, and writes both Markdown and JSON reports to the outputs directory.

### Programmatic Python API

Embed the search pipeline in Python applications:

```python
from skills.patent-search.tools.cnipa_search import main

# Search for patents by inventor "王五" with strict pagination

exit_code = main([
    "--inventor", "王五",
    "--max-pages", "2"
])
print(f"Search finished with exit code {exit_code}")

```

### Processing Generated Reports

Consume the JSON output for custom analytics or visualization:

```python
import json
from pathlib import Path

report_json = Path("outputs/patent-search/2024-05-01_14-23-10.json")
data = json.loads(report_json.read_text(encoding="utf-8"))

# Access deduplicated application records

for hit in data["hits"]:
    pub_count = len(hit["publication_records"])
    print(f"Application {hit['application_number']} – {pub_count} publications")

```

## Key Source Files

| File | Primary Function |
|------|------------------|
| [`skills/patent-search/tools/cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_search.py) | CLI orchestrator; query building, `filter_hits` deduplication |
| [`skills/patent-search/tools/cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_crawler.py) | Playwright browser automation for CNIPA navigation |
| [`skills/patent-search/tools/cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_parse.py) | HTML extraction and `EpubSearchHit` instantiation |
| [`skills/patent-search/tools/emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/emit_search_report.py) | Markdown/JSON report generation with metadata |
| [`skills/patent-disclosure/tools/crawl/cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/crawl/cnipa_epub_search.py) | Disclosure skill integration for document production |

## Summary

- The skill manages the **Chinese patent lifecycle** through a six-stage pipeline: query construction, Playwright crawling, HTML parsing, deduplication, report emission, and disclosure integration.
- **Playwright** in [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py) handles JavaScript-rendered CNIPA content unreachable by static requests.
- The **`filter_hits`** function aggregates multiple publications under single `application_number` entries to eliminate redundancy.
- Reports generate in **dual formats** (Markdown for humans, JSON for machines) with comprehensive metadata and pagination statistics.
- Modular architecture allows the `patent-disclosure` skill to consume search outputs and produce formal DOCX disclosure documents.

## Frequently Asked Questions

### How does the skill handle CNIPA's JavaScript-heavy interface?

The implementation uses **Playwright** to drive a headless browser. The [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py) module renders the CNIPA Advanced Search page dynamically, fills interactive form fields, and extracts HTML after JavaScript execution completes. This approach succeeds where traditional `requests` or `urllib` would fail to retrieve the dynamic result tables.

### What prevents duplicate patents from appearing in reports?

The `filter_hits` function normalizes identity strings and groups records by **application number**. Instead of listing each publication date separately, it creates a single entry per invention containing an aggregated `publication_records` array. This deduplication ensures that continuation publications or correction notices do not create redundant entries in the final output.

### Can searches be restricted by patent type or result volume?

Yes. The CLI accepts a `--type` argument (invention, utility model, or design) to filter patent categories. Pagination is controlled via `--max-pages` for budgeted searches, while the `--complete` flag forces comprehensive retrieval when the total result count is known and full data capture is required.

### Where are search reports stored and how are they structured?

The [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py) module writes timestamped files to `outputs/patent-search/` in two formats: **Markdown** for human review and **JSON** for programmatic consumption. Both include query parameters, execution timestamps, completeness indicators, and deduplicated hit arrays, with status summaries printed to `stderr` during execution.