# How to Perform CNIPA Bibliographic Searches Using the Patent‑Search Skill

> Learn to perform CNIPA bibliographic searches with the patent-search skill. This tool crawls the CNIPA EPUB database, filters results, and generates a Markdown report. Discover efficient patent searching.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-06

---

**The CNIPA bibliographic search is implemented as a modular command‑line tool in the patent‑search skill that uses Playwright to crawl the CNIPA EPUB database, parses results into structured data, filters by inventor or applicant, and emits a Markdown report.**

The patent‑search skill, part of the `handsomestWei/patent-disclosure-skill` repository, provides a dedicated toolchain for automating **CNIPA bibliographic searches** against the China National Intellectual Property Administration’s EPUB database. It combines headless browser automation with configurable filtering to produce structured patent reports without manual web navigation.

## Architecture Overview

The CNIPA search toolchain consists of five decoupled components designed for testability and clear separation of concerns:

- **CLI driver** ([`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)): Parses user arguments, loads configuration, and orchestrates the workflow.
- **Crawler** ([`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py)): Manages a headless Chromium instance via Playwright to submit queries and paginate results.
- **Parser** ([`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py)): Extracts structured fields—such as title, publication number, and inventor list—from the CNIPA EPUB HTML.
- **Filter & aggregator** (`filter_hits` in [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)): Applies optional inventor and applicant filters and merges duplicate publication records under the same application number.
- **Report writer** ([`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py)): Converts aggregated hits into a Markdown report saved under `outputs/patent-search/`.

The crawler can be unit‑tested in isolation using [`skills/patent-search/tests/test_cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tests/test_cnipa_search.py), and the parser remains reusable for other front‑ends.

## Step‑by‑Step Workflow

### 1. CLI Argument Handling

The entry point [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py) instantiates an `argparse.ArgumentParser` (lines 48‑77) that accepts inventor, applicant, title, IPC/LOC class, application number, publication number, patent type, and pagination controls (`--max-pages` and `--complete`).

### 2. Configuration Loading

The function `search_config.load_search_config` reads [`config.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/config.yaml) (default `max_pages: 3`) and merges command‑line overrides via `resolve_max_pages`. This determines the page budget for the crawler.

### 3. Query Construction

The helper `_query_fields` (lines 82‑95) normalizes user inputs into a dictionary passed to the crawler.

### 4. Web Crawling with Playwright

`cnipa_crawler.search_advanced` launches Chromium (`launch_chromium`), fills the advanced query form, and paginates through results. JavaScript snippets (`_RESULT_PAGE_READY_JS`, `_FETCH_RESULT_PAGE_JS`, `_FIND_NEXT_PAGE_JS`) ensure the DOM is fully loaded before extraction and detect the final page.

### 5. HTML Parsing

For each fetched page, `cnipa_parse.parse_search_result_html` transforms HTML into a list of `EpubSearchHit` objects containing `pub_number`, `application_number`, `inventors`, `applicant`, `link`, and `raw_html`.

### 6. Filtering and Aggregation

The `filter_hits` function (lines 55‑115) optionally filters by inventor identity (`normalize_identity`) and applicant aliases (`_matching_applicant`). It also consolidates multiple publication records sharing the same application number.

### 7. Completeness Tracking

The helper `completeness_note` (lines 21‑45) generates a human‑readable status indicating whether the search reached the final page, how many pages were scanned, and whether a complete crawl was requested.

### 8. Report Generation

`write_search_report` in [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py) renders the aggregated rows as a Markdown table with hyperlinks and writes the file to `outputs/patent-search/`.

## Usage Examples

### Command‑Line Searches

Run a simple inventor‑only search using default pagination:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三"

```

Search by title and IPC class, limiting to two pages:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --title "数据处理" \
    --class B01J20 \
    --max-pages 2

```

Perform a complete crawl of all pages for a specific applicant:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --applicant "华为技术有限公司" \
    --complete

```

### Programmatic Usage

You can import the underlying functions for integration into larger pipelines:

```python
from skills.patent-search.tools.cnipa_search import (
    _build_parser, filter_hits, completeness_note
)

# Build an argparse.Namespace manually

args = _build_parser().parse_args([
    "--inventor", "李四",
    "--applicant", "中科院",
    "--max-pages", "1"
])

# After crawling and parsing you obtain `hits` (list of EpubSearchHit)

filtered = filter_hits(hits, inventor=args.inventor, applicants=args.applicant)
note = completeness_note(
    complete=False,
    total_pages=5,
    pages_scanned=1,
    page_budget=args.max_pages,
    pages_remaining=4,
    has_next=True
)
print(filtered, note)

```

## Summary

- The **patent‑search skill** provides a modular command‑line interface for **CNIPA bibliographic searches** located in `skills/patent-search/tools/`.
- **Playwright** drives the crawler ([`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py)) to navigate the CNIPA EPUB site, while [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py) extracts structured `EpubSearchHit` objects.
- The `filter_hits` function deduplicates records and filters by inventor or applicant identity before report generation.
- Results are written as Markdown files to `outputs/patent-search/` via [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py).
- The tool supports both interactive CLI usage and programmatic import of its Python functions.

## Frequently Asked Questions

### What is the default pagination limit for CNIPA bibliographic searches?

The default limit is three pages, defined in [`config.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/config.yaml) and loaded by `search_config.load_search_config`. You can override this with the `--max-pages` flag or request an unlimited crawl using the `--complete` flag.

### How does the patent‑search skill handle duplicate patent records?

The `filter_hits` function in [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py) groups multiple publication records that share the same application number into a single consolidated entry, preventing redundant entries in the final report.

### Can I use the CNIPA search module programmatically without the CLI?

Yes. The functions `_build_parser`, `filter_hits`, and `completeness_note` are importable from [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py), allowing you to construct argument namespaces manually and process `EpubSearchHit` lists within Python scripts.

### What dependencies are required to run the crawler?

The crawler requires **Playwright** and a headless Chromium browser instance. The skill uses JavaScript injection (`_RESULT_PAGE_READY_JS`, `_FETCH_RESULT_PAGE_JS`) to ensure the DOM is fully loaded before parsing each result page.