How Patent-Disclosure-Skill Handles Prior Art Searching: CNIPA Integration and Fallback Logic
The patent-disclosure-skill implements a deterministic two-stage prior art search that prioritizes CNIPA EPUB data, enforces classification-based filtering, and provides a single-attempt fallback to Google Patents when primary sources fail.
This open-source skill automates the discovery and validation of prior art for Chinese patent disclosures through a structured retrieval pipeline. According to the handsomestWei/patent-disclosure-skill repository, the system integrates prior art searching as Step 5 of the disclosure workflow, combining deterministic web scraping with intelligent deduplication and strict quality controls.
The Two-Stage Prior Art Search Architecture
The skill organizes retrieval into distinct layers that progressively refine raw patent data into citation-ready entries.
Stage 1: CNIPA EPUB Retrieval via Playwright
The primary retrieval layer targets the China National Intellectual Property Administration (CNIPA) EPUB database. The driver script skills/patent-disclosure/tools/crawl/cnipa_epub_search.py executes a browser automation sequence using Playwright to scrape the official CNIPA site within a single Chromium session.
Before execution, the skill probes the environment using python skills/patent-disclosure/tools/browser.py --probe to verify Playwright availability. The search accepts 2-8 keyword units extracted from the invention description and returns a structured JSON payload stored in EPUB_HITS_JSON.
# Verify browser environment
python skills/patent-disclosure/tools/browser.py --probe
# Primary search for invention patents
python skills/patent-disclosure/tools/crawl/cnipa_epub_search.py \
--type invention \
知识库 检索增强 大语言模型
Stage 2: Classification-Based Enrichment and Filtering
The enrichment layer performs a two-round retrieval strategy defined in skills/patent-disclosure/prompts/prior_art_search.md. First, the "recall" round extracts raw hits and pulls IPC (International Patent Classification) and LOC (Locarno Classification) codes from the JSON fields ipc_codes and loc_codes.
Second, a "focused" round re-executes the search using the --class parameter with the extracted classification codes to obtain a high-relevance subset. The system merges results from both rounds on pub_number or link fields to ensure uniqueness.
# Second-round search using specific IPC classes
python skills/patent-disclosure/tools/crawl/cnipa_epub_search.py \
--type invention \
--class B01J20,B01D53 \
胺功能化
Fallback Mechanism for Insufficient Results
When the primary CNIPA channel fails—whether through non-zero exit codes, timeouts, or empty EPUB_HITS_JSON arrays—the skill downgrades to a web-search fallback. This secondary path queries Google Patents and optionally Google Scholar through the utility function google_patents_websearch_query located in skills/patent-disclosure/tools/patent_type.py.
The fallback operates under strict constraints: it permits only a single attempt and does not abort the workflow upon failure. Instead, the skill records the scarcity of hits and proceeds with available data, ensuring the disclosure pipeline remains unblocked.
from skills.patent-disclosure.tools.patent_type import google_patents_websearch_query
# Generate fallback query for Google Patents
query = google_patents_websearch_query(
core_word="胺功能化",
pat_type="invention",
class_codes="B01J20"
)
Quality Enforcement Rules
The skill enforces mandatory content policies before any prior art enters the final disclosure document.
Abstract-Driven Understanding Requirements
For every CNIPA-derived entry containing a non-empty abstract field, the skill must read and internalize the abstract content before summarizing the invention. This rule, specified in skills/patent-disclosure/prompts/prior_art_search.md, explicitly prohibits verbatim copying and requires synthesis of the technical concepts.
Link Integrity and Verification
Every cited prior-art entry must include a verifiable URL taken directly from the JSON link field. The system explicitly forbids fabricated or altered URLs, ensuring all citations remain externally verifiable.
Execution Workflow Reference
The complete prior art searching sequence executes as Step 5 in the skill pipeline, as referenced in skills/patent-disclosure/SKILL.md. The workflow follows this deterministic order:
- Environment probe via
browser.py --probe - Primary CNIPA search with keyword extraction and
--typespecification (invention, utility_model, or design) - Classification extraction and second-round focused search
- Deduplication on
pub_numberand merging of result sets - Fallback evaluation if hit counts are insufficient
- Population of the "1.1 现有技术" section with validated entries
Summary
- Primary source: CNIPA EPUB accessed via
cnipa_epub_search.pywith Playwright automation - Two-stage filtering: Initial keyword recall followed by IPC/LOC classification refinement
- Deduplication strategy: Merge results on
pub_numberandlinkfields to ensure uniqueness - Fallback protocol: Single-attempt downgrade to Google Patents via
google_patents_websearch_querywhen CNIPA fails - Quality gates: Mandatory abstract comprehension and strict URL integrity enforcement
- Integration point: Step 5 of the disclosure workflow defined in
SKILL.md
Frequently Asked Questions
What happens if the CNIPA website is unreachable during prior art searching?
If cnipa_epub_search.py exits with a non-zero code, times out, or returns empty EPUB_HITS_JSON, the skill automatically triggers the fallback mechanism. It generates a Google Patents query using google_patents_websearch_query and attempts retrieval once. If the fallback also fails, the skill logs the deficiency and continues with available data rather than blocking the disclosure workflow.
How does the skill prevent duplicate prior art entries?
The system deduplicates results by merging datasets on the pub_number field (or link as secondary key) after performing the two-round retrieval strategy. This ensures that patents retrieved via keyword search do not duplicate those found through classification code refinement.
Can the skill search for design patents as well as invention patents?
Yes. The cnipa_epub_search.py script accepts a --type parameter that supports three values: invention, utility_model, and design. This allows the prior art searching pipeline to handle Chinese design patents (外观) using the same Playwright-based retrieval logic applied to technical patents.
Why does the skill require reading abstracts before writing the disclosure?
The rule specified in skills/patent-disclosure/prompts/prior_art_search.md mandates that the system must "read and internalize" the abstract when the abstract field is non-empty. This prevents verbatim copying of patent text and ensures the disclosure contains original analysis of the prior art's technical contribution, maintaining document integrity and avoiding plagiarism.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →