# How the CNIPA Portfolio Search Navigates All Results and Filters by Applicant for Same-Name Inventors

> Learn how the CNIPA portfolio search handles inventor name disambiguation. Discover its two stage pipeline for navigating results and filtering by applicant name.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-02

---

**The CNIPA portfolio search uses a two-stage pipeline: first, `search_epub_inventor_all_pages` crawls every result page via Playwright; second, `filter_portfolio_hits` normalizes names and applies substring matching against applicant aliases to disambiguate same-name inventors.**

The `patent-disclosure-skill` repository implements a robust patent portfolio crawler for the China National Intellectual Property Administration (CNIPA) e-publication database. When dealing with common inventor names—where multiple applicants may share identically-named inventors—the system must traverse complete result sets while precisely filtering by applicant. This article explains how the implementation achieves both goals using dedicated modules in `tools/crawl/`.

## Two-Stage Architecture for Complete Result Navigation

The CNIPA portfolio search operates in distinct **crawl** and **filter** stages. This separation allows the crawler to focus purely on pagination while the filtering layer handles identity disambiguation.

### Stage 1: Crawling All Pages with Playwright

The `search_epub_inventor_all_pages` function in [`tools/crawl/cnipa_epub_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_crawler.py) drives the complete result traversal. It launches a headless Chromium session, submits the inventor query, and follows every "Next →" link until exhaustion or a user-defined limit.

```python

# Core pagination loop (simplified structure)

hits = []
while True:
    page_hits = _extract_hits_from_page(page)
    hits.extend(page_hits)
    
    next_button = page.locator("//a[contains(text(),'Next')]")
    if not next_button.count() or not next_button.is_visible():
        search.complete = True
        break
        
    if max_pages and pages_scanned >= max_pages:
        search.stop_reason = "max_pages"
        break
        
    next_button.click()
    pages_scanned += 1

```

The function returns a `SearchResult` dataclass containing:

- `hits`: List of `EpubSearchHit` objects extracted from all scanned pages
- `complete`: Boolean indicating whether every result page was visited
- `stop_reason`: `"max_pages"` if truncated, otherwise `None`
- `pages_scanned`: Actual count of pages processed
- `total_reported`: Declared result count from CNIPA's pagination display

### Stage 2: Applicant-Aware Filtering and Merging

The `filter_portfolio_hits` function in [`tools/crawl/cnipa_epub_portfolio.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_portfolio.py) processes the raw crawl results. It performs three critical operations: name normalization, exact inventor verification, and applicant substring matching.

## How Same-Name Inventor Disambiguation Works

When multiple applicants employ inventors with identical names—common in Chinese patent data—the system must ensure portfolio accuracy without false positives.

### Name Normalization Pipeline

Both inventor names and applicant aliases undergo `normalize_identity` processing:

1. Unicode NFKC normalization (collapses compatibility characters)
2. Whitespace stripping
3. Case folding to lowercase

```python
def normalize_identity(s: str) -> str:
    return unicodedata.normalize('NFKC', s).strip().lower()

```

This ensures that "张三", "张 三", and "張三" are treated as identical for matching purposes.

### Exact Inventor Verification

Each hit's `inventors` field is checked to confirm the queried inventor appears exactly. This prevents matching patents where the name appears only in applicant or title fields.

### Substring-Based Applicant Matching

The `_matching_applicant` function implements flexible alias detection:

```python
def _matching_applicant(actual: str, aliases: List[str]) -> Optional[str]:
    actual_norm = normalize_identity(actual)
    for alias in aliases:
        if normalize_identity(alias) in actual_norm:
            return alias  # Returns matching alias for attribution

    return None

```

A hit passes filtering if:

- No applicant aliases were supplied (returns all verified inventor matches)
- OR the normalized actual applicant contains any normalized alias as substring

This handles variations like "北京大学" matching "北京大学科技开发部".

### Application-Level Deduplication

CNIPA publishes multiple documents per application (pre-grant, granted, corrected). The filter merges these using a dictionary keyed by `application_number`:

```python
portfolio = {}  # application_number -> merged record

for hit in filtered_hits:
    key = normalize_identity(hit.application_number)
    if key in portfolio:
        portfolio[key]['publication_records'].append(hit.publication_record)
    else:
        portfolio[key] = create_new_entry(hit)

```

## Command-Line Usage Example

Run a complete portfolio search with applicant disambiguation:

```bash
python tools/crawl/cnipa_epub_portfolio.py \
    --inventor "张三" \
    --applicant "北京大学" \
    --applicant "清华大学" \
    --type invention \
    --max-pages 200

```

Output format:

```json
EPUB_PORTFOLIO_JSON:{
  "complete": true,
  "matched_count": 47,
  "hits": [
    {
      "matched_applicant": "北京大学",
      "identity_status": "verified",
      "application_number": "CN202310123456.7",
      "publication_records": [
        {"pub_no": "CN116000001A", "pub_date": "2023-05-01", "type": "A"},
        {"pub_no": "CN116000001B", "pub_date": "2024-02-15", "type": "B"}
      ]
    }
  ]
}

```

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`tools/crawl/cnipa_epub_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_crawler.py) | Playwright pagination, `search_epub_inventor_all_pages` |
| [`tools/crawl/cnipa_epub_portfolio.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_portfolio.py) | CLI orchestration, `filter_portfolio_hits` |
| [`tools/crawl/cnipa_epub_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_parse.py) | `EpubSearchHit` dataclass and page parsing |
| [`tools/shared/patent_type.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/shared/patent_type.py) | Patent type constants and normalization |

## Summary

- **`search_epub_inventor_all_pages`** (in [`cnipa_epub_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_epub_crawler.py)) handles complete CNIPA result traversal via Playwright, respecting `--max-pages` limits and reporting pagination status.
- **`filter_portfolio_hits`** (in [`cnipa_epub_portfolio.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_epub_portfolio.py)) applies Unicode-normalized identity verification and substring-based applicant matching to resolve same-name inventor ambiguity.
- The **substring matching strategy** allows flexible alias detection while the **application-level merge** eliminates duplicate publication entries.
- All normalization uses **NFKC Unicode standard** with case folding for cross-platform consistency.

## Frequently Asked Questions

### How does the crawler handle CNIPA's JavaScript-rendered pagination?

The implementation uses **Playwright's Chromium engine** to execute the full browser context, allowing interaction with dynamic "Next →" buttons that would be invisible to static HTML parsers. The `_extract_hits_from_page` helper parses the rendered DOM after each navigation.

### What happens if the inventor name matches but the applicant filter excludes all hits?

The `filter_portfolio_hits` function returns an empty hits array with `matched_count: 0`. This is not treated as an error—the CLI simply outputs the JSON payload with zero results, distinguishing between "no patents found" (crawler returned empty) and "no patents matched your applicant criteria" (filter removed all candidates).

### Can the applicant filter use regular expressions or exact matching only?

The current implementation uses **substring matching via `in` operator** on normalized strings. Exact matching can be approximated by providing aliases that are themselves complete applicant names. The codebase does not expose regex functionality, though the `_matching_applicant` function could be extended to support `re.search` patterns.

### How are incomplete crawls indicated in the output?

When `--max-pages` truncates pagination before exhaustion, `complete` is set to `false` and `stop_reason` contains `"max_pages"`. Consumers should treat this as a partial result warning—subsequent runs with higher limits or date-range constraints may be needed for full coverage.