How to Perform a CNIPA Prior-Art Search with Two-Stage IPC/LOC Classification
The CNIPA prior-art search uses a two-stage workflow: first retrieve broad results via free-text keywords, then filter by high-frequency IPC or LOC classification codes to surface the most relevant patents.
In the handsomestWei/patent-disclosure-skill repository, this process is implemented as an automated pipeline for generating patent disclosure documents. The two-stage IPC/LOC classification approach ensures comprehensive recall while maintaining precision, which is critical for satisfying CNIPA's disclosure requirements in section 1.1 ("Existing Technology").
First-Stage Recall: Broad Keyword Search
The initial stage casts a wide net using the CNIPA EPUB homepage's free-text query capability.
Search Parameters
- Input: 2–8 search units (keywords describing the technology)
- Output: JSON array containing patent metadata including
ipc_codesandloc_codesfor each hit - Implementation:
search_epub_keywordsfunction in [tools/crawl/cnipa_epub_search.py](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_search.py)
python tools/crawl/cnipa_epub_search.py \
--type invention \
机器学习 模型 优化
This command queries CNIPA's invention or utility model database. For design patents (外观), use --type design instead.
Output Format
The script prints a single JSON line to stdout:
EPUB_HITS_JSON: [{"pub_number":"CN119781913A","title":"...","abstract":"...","ipc_codes":["B01J20"],"link":"http://epub.cnipa.gov.cn/patent/CN119781913A"}, ...]
Diagnostic markers (e.g., EPUB_NOTE, EPUB_MERGE, EPUB_CLASS_HINT) are written to stderr for the calling agent to parse.
Second-Stage Classification: IPC/LOC Code Filtering
After analyzing first-stage results, the workflow extracts classification codes and reruns the search with tighter constraints.
Code Extraction Rules
From the JSON array, identify:
- 1–3 high-frequency IPC prefixes (e.g.,
B01J20,B01D53) for inventions and utility models - LOC numbers (e.g.,
26-05) for design patents
If no classification codes can be derived, the second stage is skipped and results are filtered by the original keywords only. This logic is defined in [prompts/disclosure/prior_art_search.md](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/prompts/disclosure/prior_art_search.md).
Advanced Search with Classification Codes
The advanced-search endpoint combines classification codes with keywords via the --class flag:
# Invention search with IPC codes
python tools/crawl/cnipa_epub_search.py \
--type invention \
--class B01J20,B01D53 \
胺功能化
# Design search with LOC code
python tools/crawl/cnipa_epub_search.py \
--type design \
--class 26-05 \
台灯
The underlying Playwright driver in [tools/crawl/cnipa_epub_crawler.py](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_crawler.py) handles both the basic homepage query and the classification-constrained advanced search.
Hit Consolidation and Back-Filling Strategy
The workflow applies specific rules to determine which patents appear in the final disclosure document.
Decision Logic
| Condition | Action |
|---|---|
| Second-stage results ≥ 4 items | Use only second-stage list (already IPC/LOC-filtered) |
| Second-stage results < 4 items | Back-fill from first-stage using backfill_hits_for_disclosure helper |
| Still insufficient after back-fill | Optionally retry with fewer keywords or neighboring IPC/LOC codes |
Critical constraint: Never fabricate entries. The backfill_hits_for_disclosure function selects patents sharing the same IPC/LOC with overlapping technical means until 4–6 patents are collected.
Optional Google Patents Fallback
When CNIPA EPUB search fails—due to network issues, zero hits, or insufficient results—the workflow may supplement with a single Google Patents query via patent_type.google_patents_websearch_query. This fallback is optional and never replaces a valid CNIPA link.
Final Output Requirements
Each selected patent is recorded in section 1.1 (Existing Technology) with:
- Publication number, title, and abstract (must be read before writing)
- Concise technical summary in the author's own words
- Direct link to CNIPA EPUB page (or Google Patents URL for fallback items only)
The abstract serves as the mandatory factual basis. Copying titles or URLs without consulting the abstract is explicitly disallowed per the prompt specifications.
Key Implementation Files
Summary
- Two-stage IPC/LOC classification guarantees relevant prior-art surfacing while maintaining a safety net of broader results
- First stage uses 2–8 free-text keywords via
cnipa_epub_search.pyto establish recall - Second stage applies 1–3 IPC prefixes or LOC codes via the advanced-search endpoint for precision
- Back-filling logic ensures 4–6 patents minimum without fabrication
- Google Patents fallback handles CNIPA service failures as optional supplement
- All outputs require abstract-verified technical summaries with direct CNIPA links
Frequently Asked Questions
What are IPC and LOC classification codes?
IPC (International Patent Classification) codes categorize inventions and utility models by technical field (e.g., B01J20 for chemical processes). LOC (洛迦诺/Locarno) codes classify industrial designs by product type (e.g., 26-05 for lamps). The CNIPA prior-art search uses these codes to filter patents sharing the same technical classification as the target invention.
How many keywords should I provide for the first-stage search?
Provide 2–8 search units (keywords or short phrases) that describe the core technology. Fewer than 2 may miss relevant prior art; more than 8 can dilute precision. The cnipa_epub_search.py script processes these as a space-separated argument list.
Can I skip the second-stage classification search?
Yes, but only when first-stage results contain no extractable IPC or LOC codes. In this case, the workflow filters first-stage results by the original keywords. However, skipping classification generally reduces precision, so the system prefers to proceed with both stages when codes are available.
What happens if the CNIPA EPUB site is unavailable?
The workflow implements a Google Patents fallback via patent_type.google_patents_websearch_query. This generates a single supplementary query but never replaces CNIPA links in the final disclosure. The calling agent detects failure conditions (network errors, zero hits, insufficient results) through stderr markers like EPUB_NOTE.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →