# De-Identification Rules for Patent Disclosures: A Technical Guide to the handsomestWei Implementation

> Learn handsomestWei patent disclosure de-identification rules. Discover how this skill uses regex and policy to safeguard PII and commercial secrets, preserving technical narrative.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: technical-guide
- Published: 2026-09-04

---

**The handsomestWei/patent-disclosure-skill repository implements a dual-layer de-identification system that combines deterministic regex-based redaction with policy-driven abstraction guidelines to remove personally identifiable information (PII) and commercial secrets from patent texts while preserving technical narrative integrity.**

Managing sensitive information in patent disclosures requires precise **de-identification rules** that balance data privacy with technical clarity. The `handsomestWei/patent-disclosure-skill` open-source project defines a comprehensive framework for patent disclosure de-identification, implemented across Python utility functions and markdown policy documentation. This technical guide examines the specific redaction patterns, placeholder conventions, and audit mechanisms that ensure compliant data sanitization.

## Core De-Identification Architecture

The repository employs a two-tier approach to sensitive data removal. First, a programmatic regex engine handles structured identifiers like phone numbers and application IDs. Second, a policy table guides abstract rewriting of business context and categorical data.

### Regex-Based Entity Redaction in redact.py

Located at [`skills/patent-oa/tools/redact.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tools/redact.py), the core `redact_text()` function applies deterministic pattern matching to six distinct entity types:

- **Company/Applicant Names**: Chinese corporate entities ending in common suffixes (科技有限公司, 有限公司, etc.) are replaced with the placeholder **"申请人甲"** (Applicant A).
- **Patent Application Numbers**: CN-prefixed application numbers transform into the generic pattern **"CNXXXXXXXXXX.X"**, with optional SHA-256 hashing available via the `hash_app_nos` parameter.
- **Mobile Phone Numbers**: 11-digit Chinese mobile formats become **"【电话已脱敏】"** (Phone redacted).
- **Email Addresses**: All email patterns substitute to **"redacted@example.com"**.
- **Identity Card Numbers**: Chinese ID numbers (18 digits) are masked as **"【证件号已脱敏】"** (ID number redacted).
- **Custom Entity Names**: Additional sensitive names pass through the `extra_names` parameter and normalize to **"【脱敏姓名】"** (Redacted name).

### Policy-Driven Abstraction Guidelines

Beyond regex matching, [`skills/patent-disclosure/prompts/invention/disclosure_builder.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/prompts/invention/disclosure_builder.md) (Section 7.5) establishes semantic **de-identification rules** for unstructured business content:

- **Industry Descriptions**: Specific business domains (e.g., "XX检测") abstract to generic scenarios like "多标签分类场景" (multi-label classification scenario).
- **Categorical Labels**: Specific category enumerations replace with generic alphabetic markers (**A**, **B**, **C**) instead of numbered lists.
- **Quantitative Values**: Precise numeric metrics transform into qualitative ranges (e.g., "每日N人" becomes "每日一定规模" meaning "daily scale of certain size").
- **Product Identifiers**: Specific system or product names substitute with "某系统" (a certain system) or complete removal.

## Technical Implementation Details

### The redact_text Function Signature

The primary de-identification interface accepts three parameters:

```python
def redact_text(text: str, extra_names: list = None, hash_app_nos: bool = False) -> tuple:
    """
    Returns: (redacted_text: str, change_log: list)
    """

```

Setting `hash_app_nos=True` replaces application numbers with SHA-256 hashes rather than the standard X-pattern, providing additional security for highly sensitive filings.

### Change-Log Generation for Audit Trails

Every redaction operation appends structured metadata to a `notes` list, enabling forensic auditing. According to the source code in [`redact.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/redact.py), the system tracks replacement counts per entity type (e.g., `'company_replacements=1'`, `'app_no_count=1'`) and maintains mapping records like `'company→申请人甲:上海某某科技有限公司'`. This audit trail persists through the ingestion pipeline defined in [`skills/patent-oa/tools/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tools/ingest_case.py), which invokes `redact_text()` before vector store insertion to ensure automated sanitization.

## Practical Code Example

The following demonstration processes a raw patent disclosure containing multiple sensitive identifiers:

```python
from redact import redact_text

# Raw disclosure containing PII

raw_text = """
上海某某科技有限公司
电话13812345678
邮箱john.doe@example.com
申请号 CN202212345678.9
身份证 310101199001011234
"""

# Execute de-identification

redacted_output, audit_notes = redact_text(raw_text)

print(redacted_output)

# Output:

# 申请人甲

# 【电话已脱敏】

# redacted@example.com

# CNXXXXXXXXXX.X

# 【证件号已脱敏】

print(audit_notes)

# Audit trail:

# ['company→申请人甲:上海某某科技有限公司', 

#  'company_replacements=1',

#  'app_no_redacted', 

#  'app_no_count=1', 

#  'phone=1', 

#  'email=1', 

#  'id=1']

```

## Summary

- **Dual-layer protection**: The system combines deterministic regex patterns in [`redact.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/redact.py) with semantic abstraction policies in [`disclosure_builder.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/disclosure_builder.md) to handle both structured identifiers and unstructured business context.
- **Deterministic placeholders**: Fixed replacement strings like "申请人甲" and "CNXXXXXXXXXX.X" ensure consistency across documents while preventing data reconstruction.
- **Comprehensive entity coverage**: The framework handles corporate names, patent numbers, contact information, government IDs, and custom entities via the `extra_names` parameter.
- **Auditable transformations**: Every redaction generates structured change logs, supporting compliance requirements and quality assurance workflows.
- **Integration ready**: The `redact_text()` function integrates directly into ingestion pipelines, as demonstrated in [`skills/patent-oa/tools/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tools/ingest_case.py), ensuring automated sanitization before vector storage.

## Frequently Asked Questions

### What types of sensitive data does the patent disclosure de-identification system handle?

The system processes six primary categories of sensitive information: Chinese company names with corporate suffixes, CN-format patent application numbers, Chinese mobile phone numbers, email addresses, Chinese national ID numbers, and custom entities provided through the `extra_names` parameter. Additionally, the policy documentation in [`disclosure_builder.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/disclosure_builder.md) covers business domain descriptions, categorical labels, and quantitative metrics that require semantic abstraction rather than simple string replacement.

### How does the system ensure that de-identified patent disclosures remain technically coherent?

Rather than deleting sensitive content entirely, the implementation uses contextual placeholders that preserve grammatical structure and technical meaning. For example, company names become "申请人甲" (Applicant A), maintaining the semantic role of the entity in the sentence, while application numbers retain the CN prefix with masked digits to preserve format recognition. The **de-identification rules** specifically mandate that "the title's domain object must stay intact" while only surrounding identifying text undergoes redaction.

### Can the de-identification rules be customized for specific patent portfolios?

Yes, the `redact_text()` function in [`skills/patent-oa/tools/redact.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tools/redact.py) accepts an `extra_names` parameter accepting a list of strings, allowing users to specify additional sensitive terms or internal codenames unique to their organization. Furthermore, the `hash_app_nos` boolean flag toggles between standardized placeholder masking (CNXXXXXXXXXX.X) and cryptographic hashing for application numbers, accommodating varying security requirements across different deployment scenarios.

### Where is the de-identification logic integrated into the patent processing pipeline?

The redaction function operates at the ingestion boundary within [`skills/patent-oa/tools/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tools/ingest_case.py), ensuring that all patent disclosures undergo sanitization before entering the knowledge base or vector store. The test suite in [`skills/patent-oa/tests/test_oa_store.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-oa/tests/test_oa_store.py) validates that these transformations produce expected placeholder outputs, maintaining data quality standards throughout the storage workflow.