CNIPA epub Search vs Patent-Search Skill for Disclosure Novelty: Key Differences Explained

The patent-search skill adds pagination safety, hit deduplication, and formatted reporting on top of the raw CNIPA epub.cnipa.gov.cn crawler to support reliable patent disclosure novelty checks.

Both approaches query the same Chinese National Intellectual Property Administration (CNIPA) database at epub.cnipa.gov.cn, but they serve different purposes within the handsomestWei/patent-disclosure-skill repository. The direct crawler provides raw access to patent data, while the skill layer transforms that data into disclosure-ready intelligence with guardrails against incomplete results.

Architecture and Entry Points

The repository separates concerns into two distinct layers. Understanding this separation clarifies which tool fits your specific novelty search requirements.

Low-Level Crawler (cnipa_crawler.py)

At the foundation sits skills/patent-search/tools/cnipa_crawler.py, a Playwright-based scraper that submits queries directly to the CNIPA advanced search form. This module fetches raw HTML pages without applying business logic, pagination caps, or result formatting. It returns unparsed page content that requires manual processing to extract publication numbers, dates, and applicant information.

High-Level Skill Interface (cnipa_search.py)

Built atop the crawler, skills/patent-search/tools/cnipa_search.py functions as the orchestration CLI. It coordinates the crawler invocation, applies configuration limits from config.yaml, and delegates formatting to emit_search_report.py. This entry point is designed specifically for conversational agents performing disclosure novelty checks, where result completeness and clarity carry legal significance.

Critical Differences in Query Execution

Pagination Control and Safety Limits

The raw crawler imposes no intrinsic limits on page retrieval, allowing unbounded requests that risk service denial or incomplete manual scans.

Conversely, the patent-search skill enforces strict pagination policies defined in config.yaml. By default, it limits searches to 3 pages with a hard ceiling of 20 pages (see lines 13‑19 of cnipa_search.py). Users must explicitly pass --max-pages or --complete flags to exceed these defaults, preventing accidental exhaustive crawling that could violate platform terms or misrepresent result completeness (lines 21‑27).

Result Aggregation and Deduplication

Direct CNIPA searches return individual publication announcements as separate entries, forcing users to manually identify when multiple publications belong to the same patent application.

The skill implements intelligent merging through the filter_hits logic (lines 55‑95 of cnipa_search.py). This function collapses multiple announcements—such as publication and grant notices—into unified records while preserving metadata like identity_status and matched_applicant. This deduplication prevents disclosure reports from listing the same invention multiple times under different publication stages.

Output Formatting and Reporting

When using the crawler directly, you receive raw HTML or extracted text requiring further processing to generate human-readable reports.

The skill automatically generates structured Markdown documents in outputs/patent-search/, timestamped with identifiers like SEARCH‑20231101‑101530.md. These reports include machine-readable prefixes (EPUB_SEARCH_MD:, EPUB_SEARCH_JSON:) for downstream agent consumption (lines 33‑36 of cnipa_search.py), while presenting human-friendly tables that explicitly note whether pagination completed or truncated results.

Practical Usage Examples

Direct Low-Level Crawler (Development Use)

Use this approach only when building custom pipelines or debugging the scraping logic itself:

python skills/patent-search/tools/cnipa_crawler.py \
    --keyword "机器学习" \
    --type invention \
    --max-pages 10

This command returns raw HTML for 10 result pages without deduplication or safety warnings about result completeness.

Patent-Search Skill (Disclosure Novelty Checks)

Use this approach for production novelty searches where result integrity matters:

python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三" \
    --applicant "华为技术有限公司" \
    --title "数据处理" \
    --class B01J20 \
    --max-pages 5

This execution respects the configured pagination limits, merges duplicate application records via filter_hits, and writes a formatted report to outputs/patent-search/ with clear pagination status indicators.

When to Use Each Approach

Choose the direct CNIPA crawler when you need unfiltered access to the HTML structure, are developing new parsing logic in cnipa_parse.py, or require result counts exceeding the skill's safety limits for research purposes.

Choose the patent-search skill when performing disclosure novelty analysis where incomplete results could create legal risk. The skill's explicit completeness warnings, deduplicated hit lists, and standardized Markdown reports align with the repository's goal of supporting transparent, auditable patent disclosures.

Summary

  • Same target, different layers: Both tools query epub.cnipa.gov.cn, but the skill adds a governance layer atop the raw crawler.
  • Safety by default: The skill enforces pagination caps (3 soft/20 hard) unless explicitly overridden, while the crawler has no limits.
  • Intelligent merging: The skill's filter_hits function deduplicates multiple publications of the same application, delivering cleaner disclosure reports.
  • Structured outputs: Unlike the crawler's raw HTML, the skill produces timestamped Markdown reports with both human and machine-readable sections.
  • Completeness transparency: The skill explicitly warns users when result sets are truncated due to pagination limits, a critical feature for novelty disclosure workflows.

Frequently Asked Questions

What file handles the deduplication logic for patent applications?

The filter_hits function in skills/patent-search/tools/cnipa_search.py (lines 55‑95) implements the deduplication logic. It merges multiple announcement records belonging to the same application number while preserving critical metadata like applicant identity, ensuring disclosure reports list each unique invention only once regardless of how many publication stages exist.

Can I perform an exhaustive search across all CNIPA result pages?

Yes, but only through explicit intent. The raw cnipa_crawler.py allows arbitrary page counts without warnings. Using the skill interface, you must pass the --complete flag or a specific --max-pages value exceeding the default 3-page limit (up to the hard ceiling of 20) to signal that you accept the risks and processing time of comprehensive retrieval (see lines 21‑27 of cnipa_search.py).

How does the patent-search skill indicate incomplete results?

The skill generates Markdown reports via emit_search_report.py that include explicit completeness notes within the document header. Additionally, the raw output prefixes EPUB_SEARCH_MD: and EPUB_SEARCH_JSON: contain metadata fields indicating whether pagination finished or terminated early, allowing downstream agents to alert users that additional prior art may exist beyond the fetched pages.

What is the difference between cnipa_parse.py and emit_search_report.py?

cnipa_parse.py extracts structured EpubSearchHit objects from the raw HTML returned by the crawler, focusing on data extraction. emit_search_report.py consumes these objects after filtering and formats them into the final Markdown presentation layer, adding human-readable explanations, pagination warnings, and legal disclaimers appropriate for disclosure contexts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →