Prior Art Search Strategies in Patent-Disclosure-Skill: A Technical Deep Dive

The patent-disclosure-skill repository implements a multi-layered prior art search workflow combining field-based query construction, applicant-alias matching, inventor verification, and configurable pagination to balance coverage and efficiency against the CNIPA database.

Patent practitioners and researchers need reliable tools to discover existing prior art before filing new applications. The patent-disclosure-skill open-source project, developed by handsomestWei, provides a specialized skill for searching China's National Intellectual Property Administration (CNIPA) patent database. This article examines the specific prior art search strategies implemented in the codebase and how they work together to deliver accurate, deduplicated results.

Field-Based Advanced Query Construction

The foundation of the search strategy lies in precise bibliographic querying. In skills/patent-search/tools/cnipa_search.py, the _query_fields helper (lines 82-94) constructs a dictionary of non-empty search parameters:


# From cnipa_search.py, lines 82-94

def _query_fields(inventor=None, applicant=None, title=None,
                  ipc_code=None, app_no=None, pub_no=None):
    """
    Build clean query dict for CNIPA advanced search endpoint.
    Only includes fields with actual values.
    """
    fields = {}
    if inventor:
        fields['inventor'] = inventor
    if applicant:
        fields['applicant'] = applicant
    # ... additional fields: title, ipc_code, app_no, pub_no

    return fields

This dictionary is passed to the CNIPA advanced query endpoint, enabling exact matches on:

  • Inventor name
  • Applicant/assignee name
  • Patent title keywords
  • IPC or LOC classification codes
  • Application number
  • Publication number

By filtering empty fields at the source, the tool avoids sending overly broad queries that would return excessive noise.

Inventor Verification for Result Accuracy

Raw database results often contain false positives where inventor names partially match or appear in unrelated contexts. The filter_hits function (lines 68-74) implements strict inventor verification:


# Inventor verification block from filter_hits

if inventor_query:
    normalized_inventors = [inv.strip().lower() for inv in hit.get('inventors', [])]
    if inventor_query.lower() not in normalized_inventors:
        continue  # Discard hit — inventor mismatch

Each hit's inventors list is normalized (stripped of whitespace, lowercased) and compared case-insensitively against the queried inventor name. Hits lacking the verified inventor are discarded before reaching the final report.

Applicant Alias Matching

Organizations frequently file patents under subsidiary names, historical names, or translated variants. The _matching_applicant function (lines 45-52) supports multiple applicant aliases:


# _matching_applicant implementation

def _matching_applicant(record_applicant, alias_list):
    """Check if record applicant matches any user-supplied alias."""
    normalized_record = normalize_applicant_name(record_applicant)
    for alias in alias_list:
        if normalized_record in normalize_applicant_name(alias):
            return True
    return False

Users provide aliases via CLI or chat interface. The helper normalizes both the actual applicant from the database record and each supplied alias, then checks for substring inclusion. This captures variations like "Beijing Xiaomi Mobile Software Co., Ltd." matching alias "Xiaomi".

Result Consolidation by Application Number

A single patent application often generates multiple publication records (initial publication, grant announcement, correction notices). The consolidation logic in filter_hits (lines 101-115) merges these into unified entries:

  • Records sharing the same application number are grouped
  • Multiple publication_records are collected into a list
  • The final output presents one row per application, not per announcement

This deduplication prevents practitioners from reviewing the same underlying invention multiple times and keeps search reports concise.

Configurable Pagination Control

The tool implements a two-tier pagination strategy defined in skills/patent-search/tools/search_config.py (lines 8-58):

Setting Value Purpose
default_max_pages 3 Conservative default for quick screening
max_pages_hard 20 Absolute ceiling to prevent runaway crawling

The resolve_max_pages function calculates the effective budget:


# From search_config.py, lines 8-58

DEFAULTS = {
    'max_pages': 3,
    'max_pages_hard': 20,
}

def resolve_max_pages(user_max=None, complete_mode=False):
    """
    Determine page budget based on user flags and hard limits.
    """
    if complete_mode:
        # Will expand to hard limit after total count known

        return 'complete'
    if user_max is None:
        return DEFAULTS['max_pages']
    return min(int(user_max), DEFAULTS['max_pages_hard'])

Users control pagination via:

  • --max-pages N: Explicitly request N pages (capped at 20)
  • --complete: Signal intent for exhaustive search; tool expands to hard limit after determining total result count

This design balances speed (default 3 pages) against comprehensiveness (up to 20 pages), with clear user intent signals.

Search Report Generation

After filtering, the tool assembles a structured payload and invokes emit_search_report.write_search_report. The resulting Markdown report (written to outputs/patent-search/) includes:

  • Total pages available vs. pages actually scanned
  • Completeness notation (whether search was exhaustive or partial)
  • Count of matched, verified hits
  • Detailed hit metadata with consolidated publication history

Summary

The prior art search strategies in patent-disclosure-skill deliver a controlled, verifiable workflow:

  • Field-based queries target specific bibliographic data in the CNIPA database
  • Inventor verification eliminates false positives through strict name matching
  • Applicant alias matching catches organizational name variations
  • Application-based consolidation deduplicates multiple publication events
  • Configurable pagination lets users trade speed for coverage with explicit flags
  • Structured reporting documents search scope and completeness for legal records

Frequently Asked Questions

The skill searches the CNIPA (China National Intellectual Property Administration) database through its advanced query endpoint. The crawler uses Playwright to interact with the web interface, as implemented in cnipa_crawler.search_advanced.

How does the tool prevent false positives from similar inventor names?

The filter_hits function (lines 68-74 in cnipa_search.py) normalizes inventor names from each database hit and performs case-insensitive exact matching against the queried name. Hits without a verified match are discarded before reaching the final report.

Can I search for patents filed by a company with multiple subsidiary names?

Yes. The _matching_applicant function accepts a list of aliases and normalizes both database records and your supplied names to catch variations. Pass aliases via the applicant field to capture subsidiaries, historical names, or translation differences.

What happens if I need to search more than the default 3 pages?

Use the --max-pages N flag to request a specific page count (maximum 20), or --complete to signal exhaustive search intent. The resolve_max_pages function in search_config.py enforces a hard ceiling of 20 pages to prevent excessive crawling of the CNIPA portal.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →