Prior Art Search Strategies in Patent-Disclosure-Skill: A Technical Deep Dive
The patent-disclosure-skill repository implements a multi-layered prior art search workflow combining field-based query construction, applicant-alias matching, inventor verification, and configurable pagination to balance coverage and efficiency against the CNIPA database.
Patent practitioners and researchers need reliable tools to discover existing prior art before filing new applications. The patent-disclosure-skill open-source project, developed by handsomestWei, provides a specialized skill for searching China's National Intellectual Property Administration (CNIPA) patent database. This article examines the specific prior art search strategies implemented in the codebase and how they work together to deliver accurate, deduplicated results.
Field-Based Advanced Query Construction
The foundation of the search strategy lies in precise bibliographic querying. In skills/patent-search/tools/cnipa_search.py, the _query_fields helper (lines 82-94) constructs a dictionary of non-empty search parameters:
# From cnipa_search.py, lines 82-94
def _query_fields(inventor=None, applicant=None, title=None,
ipc_code=None, app_no=None, pub_no=None):
"""
Build clean query dict for CNIPA advanced search endpoint.
Only includes fields with actual values.
"""
fields = {}
if inventor:
fields['inventor'] = inventor
if applicant:
fields['applicant'] = applicant
# ... additional fields: title, ipc_code, app_no, pub_no
return fields
This dictionary is passed to the CNIPA advanced query endpoint, enabling exact matches on:
- Inventor name
- Applicant/assignee name
- Patent title keywords
- IPC or LOC classification codes
- Application number
- Publication number
By filtering empty fields at the source, the tool avoids sending overly broad queries that would return excessive noise.
Inventor Verification for Result Accuracy
Raw database results often contain false positives where inventor names partially match or appear in unrelated contexts. The filter_hits function (lines 68-74) implements strict inventor verification:
# Inventor verification block from filter_hits
if inventor_query:
normalized_inventors = [inv.strip().lower() for inv in hit.get('inventors', [])]
if inventor_query.lower() not in normalized_inventors:
continue # Discard hit — inventor mismatch
Each hit's inventors list is normalized (stripped of whitespace, lowercased) and compared case-insensitively against the queried inventor name. Hits lacking the verified inventor are discarded before reaching the final report.
Applicant Alias Matching
Organizations frequently file patents under subsidiary names, historical names, or translated variants. The _matching_applicant function (lines 45-52) supports multiple applicant aliases:
# _matching_applicant implementation
def _matching_applicant(record_applicant, alias_list):
"""Check if record applicant matches any user-supplied alias."""
normalized_record = normalize_applicant_name(record_applicant)
for alias in alias_list:
if normalized_record in normalize_applicant_name(alias):
return True
return False
Users provide aliases via CLI or chat interface. The helper normalizes both the actual applicant from the database record and each supplied alias, then checks for substring inclusion. This captures variations like "Beijing Xiaomi Mobile Software Co., Ltd." matching alias "Xiaomi".
Result Consolidation by Application Number
A single patent application often generates multiple publication records (initial publication, grant announcement, correction notices). The consolidation logic in filter_hits (lines 101-115) merges these into unified entries:
- Records sharing the same application number are grouped
- Multiple
publication_recordsare collected into a list - The final output presents one row per application, not per announcement
This deduplication prevents practitioners from reviewing the same underlying invention multiple times and keeps search reports concise.
Configurable Pagination Control
The tool implements a two-tier pagination strategy defined in skills/patent-search/tools/search_config.py (lines 8-58):
| Setting | Value | Purpose |
|---|---|---|
default_max_pages |
3 | Conservative default for quick screening |
max_pages_hard |
20 | Absolute ceiling to prevent runaway crawling |
The resolve_max_pages function calculates the effective budget:
# From search_config.py, lines 8-58
DEFAULTS = {
'max_pages': 3,
'max_pages_hard': 20,
}
def resolve_max_pages(user_max=None, complete_mode=False):
"""
Determine page budget based on user flags and hard limits.
"""
if complete_mode:
# Will expand to hard limit after total count known
return 'complete'
if user_max is None:
return DEFAULTS['max_pages']
return min(int(user_max), DEFAULTS['max_pages_hard'])
Users control pagination via:
--max-pages N: Explicitly request N pages (capped at 20)--complete: Signal intent for exhaustive search; tool expands to hard limit after determining total result count
This design balances speed (default 3 pages) against comprehensiveness (up to 20 pages), with clear user intent signals.
Search Report Generation
After filtering, the tool assembles a structured payload and invokes emit_search_report.write_search_report. The resulting Markdown report (written to outputs/patent-search/) includes:
- Total pages available vs. pages actually scanned
- Completeness notation (whether search was exhaustive or partial)
- Count of matched, verified hits
- Detailed hit metadata with consolidated publication history
Summary
The prior art search strategies in patent-disclosure-skill deliver a controlled, verifiable workflow:
- Field-based queries target specific bibliographic data in the CNIPA database
- Inventor verification eliminates false positives through strict name matching
- Applicant alias matching catches organizational name variations
- Application-based consolidation deduplicates multiple publication events
- Configurable pagination lets users trade speed for coverage with explicit flags
- Structured reporting documents search scope and completeness for legal records
Frequently Asked Questions
What patent database does patent-disclosure-skill search?
The skill searches the CNIPA (China National Intellectual Property Administration) database through its advanced query endpoint. The crawler uses Playwright to interact with the web interface, as implemented in cnipa_crawler.search_advanced.
How does the tool prevent false positives from similar inventor names?
The filter_hits function (lines 68-74 in cnipa_search.py) normalizes inventor names from each database hit and performs case-insensitive exact matching against the queried name. Hits without a verified match are discarded before reaching the final report.
Can I search for patents filed by a company with multiple subsidiary names?
Yes. The _matching_applicant function accepts a list of aliases and normalizes both database records and your supplied names to catch variations. Pass aliases via the applicant field to capture subsidiaries, historical names, or translation differences.
What happens if I need to search more than the default 3 pages?
Use the --max-pages N flag to request a specific page count (maximum 20), or --complete to signal exhaustive search intent. The resolve_max_pages function in search_config.py enforces a hard ceiling of 20 pages to prevent excessive crawling of the CNIPA portal.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →