How the CNIPA Portfolio Search Navigates All Results and Filters by Applicant for Same-Name Inventors

The CNIPA portfolio search uses a two-stage pipeline: first, search_epub_inventor_all_pages crawls every result page via Playwright; second, filter_portfolio_hits normalizes names and applies substring matching against applicant aliases to disambiguate same-name inventors.

The patent-disclosure-skill repository implements a robust patent portfolio crawler for the China National Intellectual Property Administration (CNIPA) e-publication database. When dealing with common inventor names—where multiple applicants may share identically-named inventors—the system must traverse complete result sets while precisely filtering by applicant. This article explains how the implementation achieves both goals using dedicated modules in tools/crawl/.

Two-Stage Architecture for Complete Result Navigation

The CNIPA portfolio search operates in distinct crawl and filter stages. This separation allows the crawler to focus purely on pagination while the filtering layer handles identity disambiguation.

Stage 1: Crawling All Pages with Playwright

The search_epub_inventor_all_pages function in tools/crawl/cnipa_epub_crawler.py drives the complete result traversal. It launches a headless Chromium session, submits the inventor query, and follows every "Next →" link until exhaustion or a user-defined limit.


# Core pagination loop (simplified structure)

hits = []
while True:
    page_hits = _extract_hits_from_page(page)
    hits.extend(page_hits)
    
    next_button = page.locator("//a[contains(text(),'Next')]")
    if not next_button.count() or not next_button.is_visible():
        search.complete = True
        break
        
    if max_pages and pages_scanned >= max_pages:
        search.stop_reason = "max_pages"
        break
        
    next_button.click()
    pages_scanned += 1

The function returns a SearchResult dataclass containing:

  • hits: List of EpubSearchHit objects extracted from all scanned pages
  • complete: Boolean indicating whether every result page was visited
  • stop_reason: "max_pages" if truncated, otherwise None
  • pages_scanned: Actual count of pages processed
  • total_reported: Declared result count from CNIPA's pagination display

Stage 2: Applicant-Aware Filtering and Merging

The filter_portfolio_hits function in tools/crawl/cnipa_epub_portfolio.py processes the raw crawl results. It performs three critical operations: name normalization, exact inventor verification, and applicant substring matching.

How Same-Name Inventor Disambiguation Works

When multiple applicants employ inventors with identical names—common in Chinese patent data—the system must ensure portfolio accuracy without false positives.

Name Normalization Pipeline

Both inventor names and applicant aliases undergo normalize_identity processing:

  1. Unicode NFKC normalization (collapses compatibility characters)
  2. Whitespace stripping
  3. Case folding to lowercase
def normalize_identity(s: str) -> str:
    return unicodedata.normalize('NFKC', s).strip().lower()

This ensures that "张三", "张 三", and "張三" are treated as identical for matching purposes.

Exact Inventor Verification

Each hit's inventors field is checked to confirm the queried inventor appears exactly. This prevents matching patents where the name appears only in applicant or title fields.

Substring-Based Applicant Matching

The _matching_applicant function implements flexible alias detection:

def _matching_applicant(actual: str, aliases: List[str]) -> Optional[str]:
    actual_norm = normalize_identity(actual)
    for alias in aliases:
        if normalize_identity(alias) in actual_norm:
            return alias  # Returns matching alias for attribution

    return None

A hit passes filtering if:

  • No applicant aliases were supplied (returns all verified inventor matches)
  • OR the normalized actual applicant contains any normalized alias as substring

This handles variations like "北京大学" matching "北京大学科技开发部".

Application-Level Deduplication

CNIPA publishes multiple documents per application (pre-grant, granted, corrected). The filter merges these using a dictionary keyed by application_number:

portfolio = {}  # application_number -> merged record

for hit in filtered_hits:
    key = normalize_identity(hit.application_number)
    if key in portfolio:
        portfolio[key]['publication_records'].append(hit.publication_record)
    else:
        portfolio[key] = create_new_entry(hit)

Command-Line Usage Example

Run a complete portfolio search with applicant disambiguation:

python tools/crawl/cnipa_epub_portfolio.py \
    --inventor "张三" \
    --applicant "北京大学" \
    --applicant "清华大学" \
    --type invention \
    --max-pages 200

Output format:

EPUB_PORTFOLIO_JSON:{
  "complete": true,
  "matched_count": 47,
  "hits": [
    {
      "matched_applicant": "北京大学",
      "identity_status": "verified",
      "application_number": "CN202310123456.7",
      "publication_records": [
        {"pub_no": "CN116000001A", "pub_date": "2023-05-01", "type": "A"},
        {"pub_no": "CN116000001B", "pub_date": "2024-02-15", "type": "B"}
      ]
    }
  ]
}

Key Implementation Files

File Purpose
tools/crawl/cnipa_epub_crawler.py Playwright pagination, search_epub_inventor_all_pages
tools/crawl/cnipa_epub_portfolio.py CLI orchestration, filter_portfolio_hits
tools/crawl/cnipa_epub_parse.py EpubSearchHit dataclass and page parsing
tools/shared/patent_type.py Patent type constants and normalization

Summary

  • search_epub_inventor_all_pages (in cnipa_epub_crawler.py) handles complete CNIPA result traversal via Playwright, respecting --max-pages limits and reporting pagination status.
  • filter_portfolio_hits (in cnipa_epub_portfolio.py) applies Unicode-normalized identity verification and substring-based applicant matching to resolve same-name inventor ambiguity.
  • The substring matching strategy allows flexible alias detection while the application-level merge eliminates duplicate publication entries.
  • All normalization uses NFKC Unicode standard with case folding for cross-platform consistency.

Frequently Asked Questions

How does the crawler handle CNIPA's JavaScript-rendered pagination?

The implementation uses Playwright's Chromium engine to execute the full browser context, allowing interaction with dynamic "Next →" buttons that would be invisible to static HTML parsers. The _extract_hits_from_page helper parses the rendered DOM after each navigation.

What happens if the inventor name matches but the applicant filter excludes all hits?

The filter_portfolio_hits function returns an empty hits array with matched_count: 0. This is not treated as an error—the CLI simply outputs the JSON payload with zero results, distinguishing between "no patents found" (crawler returned empty) and "no patents matched your applicant criteria" (filter removed all candidates).

Can the applicant filter use regular expressions or exact matching only?

The current implementation uses substring matching via in operator on normalized strings. Exact matching can be approximated by providing aliases that are themselves complete applicant names. The codebase does not expose regex functionality, though the _matching_applicant function could be extended to support re.search patterns.

How are incomplete crawls indicated in the output?

When --max-pages truncates pagination before exhaustion, complete is set to false and stop_reason contains "max_pages". Consumers should treat this as a partial result warning—subsequent runs with higher limits or date-range constraints may be needed for full coverage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →