# How Patent-Disclosure-Skill Handles Prior Art Searching: CNIPA Integration and Fallback Logic

> Learn how patent-disclosure-skill performs prior art searching with CNIPA integration and fallback logic. Prioritizes CNIPA data, filters by classification, and uses Google Patents as a backup.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-04

---

**The patent-disclosure-skill implements a deterministic two-stage prior art search that prioritizes CNIPA EPUB data, enforces classification-based filtering, and provides a single-attempt fallback to Google Patents when primary sources fail.**

This open-source skill automates the discovery and validation of prior art for Chinese patent disclosures through a structured retrieval pipeline. According to the `handsomestWei/patent-disclosure-skill` repository, the system integrates **prior art searching** as Step 5 of the disclosure workflow, combining deterministic web scraping with intelligent deduplication and strict quality controls.

## The Two-Stage Prior Art Search Architecture

The skill organizes retrieval into distinct layers that progressively refine raw patent data into citation-ready entries.

### Stage 1: CNIPA EPUB Retrieval via Playwright

The primary retrieval layer targets the **China National Intellectual Property Administration (CNIPA) EPUB** database. The driver script [`skills/patent-disclosure/tools/crawl/cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/crawl/cnipa_epub_search.py) executes a browser automation sequence using Playwright to scrape the official CNIPA site within a single Chromium session.

Before execution, the skill probes the environment using `python skills/patent-disclosure/tools/browser.py --probe` to verify Playwright availability. The search accepts 2-8 keyword units extracted from the invention description and returns a structured JSON payload stored in `EPUB_HITS_JSON`.

```bash

# Verify browser environment

python skills/patent-disclosure/tools/browser.py --probe

# Primary search for invention patents

python skills/patent-disclosure/tools/crawl/cnipa_epub_search.py \
    --type invention \
    知识库 检索增强 大语言模型

```

### Stage 2: Classification-Based Enrichment and Filtering

The enrichment layer performs a **two-round retrieval strategy** defined in [`skills/patent-disclosure/prompts/prior_art_search.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/prompts/prior_art_search.md). First, the "recall" round extracts raw hits and pulls **IPC (International Patent Classification)** and **LOC (Locarno Classification)** codes from the JSON fields `ipc_codes` and `loc_codes`.

Second, a "focused" round re-executes the search using the `--class` parameter with the extracted classification codes to obtain a high-relevance subset. The system merges results from both rounds on `pub_number` or `link` fields to ensure uniqueness.

```bash

# Second-round search using specific IPC classes

python skills/patent-disclosure/tools/crawl/cnipa_epub_search.py \
    --type invention \
    --class B01J20,B01D53 \
    胺功能化

```

## Fallback Mechanism for Insufficient Results

When the primary CNIPA channel fails—whether through non-zero exit codes, timeouts, or empty `EPUB_HITS_JSON` arrays—the skill **downgrades** to a web-search fallback. This secondary path queries **Google Patents** and optionally Google Scholar through the utility function `google_patents_websearch_query` located in [`skills/patent-disclosure/tools/patent_type.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/patent_type.py).

The fallback operates under strict constraints: it permits **only a single attempt** and does not abort the workflow upon failure. Instead, the skill records the scarcity of hits and proceeds with available data, ensuring the disclosure pipeline remains unblocked.

```python
from skills.patent-disclosure.tools.patent_type import google_patents_websearch_query

# Generate fallback query for Google Patents

query = google_patents_websearch_query(
    core_word="胺功能化",
    pat_type="invention",
    class_codes="B01J20"
)

```

## Quality Enforcement Rules

The skill enforces mandatory content policies before any prior art enters the final disclosure document.

### Abstract-Driven Understanding Requirements

For every CNIPA-derived entry containing a non-empty `abstract` field, the skill **must** read and internalize the abstract content before summarizing the invention. This rule, specified in [`skills/patent-disclosure/prompts/prior_art_search.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/prompts/prior_art_search.md), explicitly prohibits verbatim copying and requires synthesis of the technical concepts.

### Link Integrity and Verification

Every cited prior-art entry must include a verifiable URL taken directly from the JSON `link` field. The system explicitly forbids fabricated or altered URLs, ensuring all citations remain externally verifiable.

## Execution Workflow Reference

The complete prior art searching sequence executes as Step 5 in the skill pipeline, as referenced in [`skills/patent-disclosure/SKILL.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/SKILL.md). The workflow follows this deterministic order:

1. **Environment probe** via `browser.py --probe`
2. **Primary CNIPA search** with keyword extraction and `--type` specification (invention, utility_model, or design)
3. **Classification extraction** and second-round focused search
4. **Deduplication** on `pub_number` and merging of result sets
5. **Fallback evaluation** if hit counts are insufficient
6. **Population** of the "1.1 现有技术" section with validated entries

## Summary

- **Primary source**: CNIPA EPUB accessed via [`cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_epub_search.py) with Playwright automation
- **Two-stage filtering**: Initial keyword recall followed by IPC/LOC classification refinement
- **Deduplication strategy**: Merge results on `pub_number` and `link` fields to ensure uniqueness
- **Fallback protocol**: Single-attempt downgrade to Google Patents via `google_patents_websearch_query` when CNIPA fails
- **Quality gates**: Mandatory abstract comprehension and strict URL integrity enforcement
- **Integration point**: Step 5 of the disclosure workflow defined in [`SKILL.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/SKILL.md)

## Frequently Asked Questions

### What happens if the CNIPA website is unreachable during prior art searching?

If [`cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_epub_search.py) exits with a non-zero code, times out, or returns empty `EPUB_HITS_JSON`, the skill automatically triggers the fallback mechanism. It generates a Google Patents query using `google_patents_websearch_query` and attempts retrieval once. If the fallback also fails, the skill logs the deficiency and continues with available data rather than blocking the disclosure workflow.

### How does the skill prevent duplicate prior art entries?

The system deduplicates results by merging datasets on the `pub_number` field (or `link` as secondary key) after performing the two-round retrieval strategy. This ensures that patents retrieved via keyword search do not duplicate those found through classification code refinement.

### Can the skill search for design patents as well as invention patents?

Yes. The [`cnipa_epub_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_epub_search.py) script accepts a `--type` parameter that supports three values: `invention`, `utility_model`, and `design`. This allows the prior art searching pipeline to handle Chinese design patents (外观) using the same Playwright-based retrieval logic applied to technical patents.

### Why does the skill require reading abstracts before writing the disclosure?

The rule specified in [`skills/patent-disclosure/prompts/prior_art_search.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/prompts/prior_art_search.md) mandates that the system must "read and internalize" the abstract when the `abstract` field is non-empty. This prevents verbatim copying of patent text and ensures the disclosure contains original analysis of the prior art's technical contribution, maintaining document integrity and avoiding plagiarism.