# CNIPA epub Search vs Patent-Search Skill for Disclosure Novelty: Key Differences Explained

> Discover the key differences between CNIPA epub search and the patent search skill for disclosure novelty. Enhance your patent analysis with added safety and reporting features.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: deep-dive
- Published: 2026-09-06

---

**The patent-search skill adds pagination safety, hit deduplication, and formatted reporting on top of the raw CNIPA epub.cnipa.gov.cn crawler to support reliable patent disclosure novelty checks.**

Both approaches query the same Chinese National Intellectual Property Administration (CNIPA) database at `epub.cnipa.gov.cn`, but they serve different purposes within the `handsomestWei/patent-disclosure-skill` repository. The direct crawler provides raw access to patent data, while the skill layer transforms that data into disclosure-ready intelligence with guardrails against incomplete results.

## Architecture and Entry Points

The repository separates concerns into two distinct layers. Understanding this separation clarifies which tool fits your specific novelty search requirements.

### Low-Level Crawler ([`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py))

At the foundation sits [`skills/patent-search/tools/cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_crawler.py), a Playwright-based scraper that submits queries directly to the CNIPA advanced search form. This module fetches raw HTML pages without applying business logic, pagination caps, or result formatting. It returns unparsed page content that requires manual processing to extract publication numbers, dates, and applicant information.

### High-Level Skill Interface ([`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py))

Built atop the crawler, [`skills/patent-search/tools/cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_search.py) functions as the orchestration CLI. It coordinates the crawler invocation, applies configuration limits from [`config.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/config.yaml), and delegates formatting to [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py). This entry point is designed specifically for conversational agents performing disclosure novelty checks, where result completeness and clarity carry legal significance.

## Critical Differences in Query Execution

### Pagination Control and Safety Limits

The raw crawler imposes no intrinsic limits on page retrieval, allowing unbounded requests that risk service denial or incomplete manual scans.

Conversely, the patent-search skill enforces strict pagination policies defined in [`config.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/config.yaml). By default, it limits searches to **3 pages** with a hard ceiling of **20 pages** (see lines 13‑19 of [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)). Users must explicitly pass `--max-pages` or `--complete` flags to exceed these defaults, preventing accidental exhaustive crawling that could violate platform terms or misrepresent result completeness (lines 21‑27).

### Result Aggregation and Deduplication

Direct CNIPA searches return individual publication announcements as separate entries, forcing users to manually identify when multiple publications belong to the same patent application.

The skill implements intelligent merging through the `filter_hits` logic (lines 55‑95 of [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)). This function collapses multiple announcements—such as publication and grant notices—into unified records while preserving metadata like `identity_status` and `matched_applicant`. This deduplication prevents disclosure reports from listing the same invention multiple times under different publication stages.

### Output Formatting and Reporting

When using the crawler directly, you receive raw HTML or extracted text requiring further processing to generate human-readable reports.

The skill automatically generates structured Markdown documents in `outputs/patent-search/`, timestamped with identifiers like `SEARCH‑20231101‑101530.md`. These reports include machine-readable prefixes (`EPUB_SEARCH_MD:`, `EPUB_SEARCH_JSON:`) for downstream agent consumption (lines 33‑36 of [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)), while presenting human-friendly tables that explicitly note whether pagination completed or truncated results.

## Practical Usage Examples

### Direct Low-Level Crawler (Development Use)

Use this approach only when building custom pipelines or debugging the scraping logic itself:

```bash
python skills/patent-search/tools/cnipa_crawler.py \
    --keyword "机器学习" \
    --type invention \
    --max-pages 10

```

This command returns raw HTML for 10 result pages without deduplication or safety warnings about result completeness.

### Patent-Search Skill (Disclosure Novelty Checks)

Use this approach for production novelty searches where result integrity matters:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三" \
    --applicant "华为技术有限公司" \
    --title "数据处理" \
    --class B01J20 \
    --max-pages 5

```

This execution respects the configured pagination limits, merges duplicate application records via `filter_hits`, and writes a formatted report to `outputs/patent-search/` with clear pagination status indicators.

## When to Use Each Approach

**Choose the direct CNIPA crawler** when you need unfiltered access to the HTML structure, are developing new parsing logic in [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py), or require result counts exceeding the skill's safety limits for research purposes.

**Choose the patent-search skill** when performing disclosure novelty analysis where incomplete results could create legal risk. The skill's explicit completeness warnings, deduplicated hit lists, and standardized Markdown reports align with the repository's goal of supporting transparent, auditable patent disclosures.

## Summary

- **Same target, different layers**: Both tools query `epub.cnipa.gov.cn`, but the skill adds a governance layer atop the raw crawler.
- **Safety by default**: The skill enforces pagination caps (3 soft/20 hard) unless explicitly overridden, while the crawler has no limits.
- **Intelligent merging**: The skill's `filter_hits` function deduplicates multiple publications of the same application, delivering cleaner disclosure reports.
- **Structured outputs**: Unlike the crawler's raw HTML, the skill produces timestamped Markdown reports with both human and machine-readable sections.
- **Completeness transparency**: The skill explicitly warns users when result sets are truncated due to pagination limits, a critical feature for novelty disclosure workflows.

## Frequently Asked Questions

### What file handles the deduplication logic for patent applications?

The `filter_hits` function in [`skills/patent-search/tools/cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_search.py) (lines 55‑95) implements the deduplication logic. It merges multiple announcement records belonging to the same application number while preserving critical metadata like applicant identity, ensuring disclosure reports list each unique invention only once regardless of how many publication stages exist.

### Can I perform an exhaustive search across all CNIPA result pages?

Yes, but only through explicit intent. The raw [`cnipa_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_crawler.py) allows arbitrary page counts without warnings. Using the skill interface, you must pass the `--complete` flag or a specific `--max-pages` value exceeding the default 3-page limit (up to the hard ceiling of 20) to signal that you accept the risks and processing time of comprehensive retrieval (see lines 21‑27 of [`cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_search.py)).

### How does the patent-search skill indicate incomplete results?

The skill generates Markdown reports via [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py) that include explicit completeness notes within the document header. Additionally, the raw output prefixes `EPUB_SEARCH_MD:` and `EPUB_SEARCH_JSON:` contain metadata fields indicating whether pagination finished or terminated early, allowing downstream agents to alert users that additional prior art may exist beyond the fetched pages.

### What is the difference between [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py) and [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py)?

[`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py) extracts structured `EpubSearchHit` objects from the raw HTML returned by the crawler, focusing on data extraction. [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py) consumes these objects after filtering and formats them into the final Markdown presentation layer, adding human-readable explanations, pagination warnings, and legal disclaimers appropriate for disclosure contexts.