De-Identification Rules for Patent Disclosures: A Technical Guide to the handsomestWei Implementation
The handsomestWei/patent-disclosure-skill repository implements a dual-layer de-identification system that combines deterministic regex-based redaction with policy-driven abstraction guidelines to remove personally identifiable information (PII) and commercial secrets from patent texts while preserving technical narrative integrity.
Managing sensitive information in patent disclosures requires precise de-identification rules that balance data privacy with technical clarity. The handsomestWei/patent-disclosure-skill open-source project defines a comprehensive framework for patent disclosure de-identification, implemented across Python utility functions and markdown policy documentation. This technical guide examines the specific redaction patterns, placeholder conventions, and audit mechanisms that ensure compliant data sanitization.
Core De-Identification Architecture
The repository employs a two-tier approach to sensitive data removal. First, a programmatic regex engine handles structured identifiers like phone numbers and application IDs. Second, a policy table guides abstract rewriting of business context and categorical data.
Regex-Based Entity Redaction in redact.py
Located at skills/patent-oa/tools/redact.py, the core redact_text() function applies deterministic pattern matching to six distinct entity types:
- Company/Applicant Names: Chinese corporate entities ending in common suffixes (科技有限公司, 有限公司, etc.) are replaced with the placeholder "申请人甲" (Applicant A).
- Patent Application Numbers: CN-prefixed application numbers transform into the generic pattern "CNXXXXXXXXXX.X", with optional SHA-256 hashing available via the
hash_app_nosparameter. - Mobile Phone Numbers: 11-digit Chinese mobile formats become "【电话已脱敏】" (Phone redacted).
- Email Addresses: All email patterns substitute to "redacted@example.com".
- Identity Card Numbers: Chinese ID numbers (18 digits) are masked as "【证件号已脱敏】" (ID number redacted).
- Custom Entity Names: Additional sensitive names pass through the
extra_namesparameter and normalize to "【脱敏姓名】" (Redacted name).
Policy-Driven Abstraction Guidelines
Beyond regex matching, skills/patent-disclosure/prompts/invention/disclosure_builder.md (Section 7.5) establishes semantic de-identification rules for unstructured business content:
- Industry Descriptions: Specific business domains (e.g., "XX检测") abstract to generic scenarios like "多标签分类场景" (multi-label classification scenario).
- Categorical Labels: Specific category enumerations replace with generic alphabetic markers (A, B, C) instead of numbered lists.
- Quantitative Values: Precise numeric metrics transform into qualitative ranges (e.g., "每日N人" becomes "每日一定规模" meaning "daily scale of certain size").
- Product Identifiers: Specific system or product names substitute with "某系统" (a certain system) or complete removal.
Technical Implementation Details
The redact_text Function Signature
The primary de-identification interface accepts three parameters:
def redact_text(text: str, extra_names: list = None, hash_app_nos: bool = False) -> tuple:
"""
Returns: (redacted_text: str, change_log: list)
"""
Setting hash_app_nos=True replaces application numbers with SHA-256 hashes rather than the standard X-pattern, providing additional security for highly sensitive filings.
Change-Log Generation for Audit Trails
Every redaction operation appends structured metadata to a notes list, enabling forensic auditing. According to the source code in redact.py, the system tracks replacement counts per entity type (e.g., 'company_replacements=1', 'app_no_count=1') and maintains mapping records like 'company→申请人甲:上海某某科技有限公司'. This audit trail persists through the ingestion pipeline defined in skills/patent-oa/tools/ingest_case.py, which invokes redact_text() before vector store insertion to ensure automated sanitization.
Practical Code Example
The following demonstration processes a raw patent disclosure containing multiple sensitive identifiers:
from redact import redact_text
# Raw disclosure containing PII
raw_text = """
上海某某科技有限公司
电话13812345678
邮箱john.doe@example.com
申请号 CN202212345678.9
身份证 310101199001011234
"""
# Execute de-identification
redacted_output, audit_notes = redact_text(raw_text)
print(redacted_output)
# Output:
# 申请人甲
# 【电话已脱敏】
# redacted@example.com
# CNXXXXXXXXXX.X
# 【证件号已脱敏】
print(audit_notes)
# Audit trail:
# ['company→申请人甲:上海某某科技有限公司',
# 'company_replacements=1',
# 'app_no_redacted',
# 'app_no_count=1',
# 'phone=1',
# 'email=1',
# 'id=1']
Summary
- Dual-layer protection: The system combines deterministic regex patterns in
redact.pywith semantic abstraction policies indisclosure_builder.mdto handle both structured identifiers and unstructured business context. - Deterministic placeholders: Fixed replacement strings like "申请人甲" and "CNXXXXXXXXXX.X" ensure consistency across documents while preventing data reconstruction.
- Comprehensive entity coverage: The framework handles corporate names, patent numbers, contact information, government IDs, and custom entities via the
extra_namesparameter. - Auditable transformations: Every redaction generates structured change logs, supporting compliance requirements and quality assurance workflows.
- Integration ready: The
redact_text()function integrates directly into ingestion pipelines, as demonstrated inskills/patent-oa/tools/ingest_case.py, ensuring automated sanitization before vector storage.
Frequently Asked Questions
What types of sensitive data does the patent disclosure de-identification system handle?
The system processes six primary categories of sensitive information: Chinese company names with corporate suffixes, CN-format patent application numbers, Chinese mobile phone numbers, email addresses, Chinese national ID numbers, and custom entities provided through the extra_names parameter. Additionally, the policy documentation in disclosure_builder.md covers business domain descriptions, categorical labels, and quantitative metrics that require semantic abstraction rather than simple string replacement.
How does the system ensure that de-identified patent disclosures remain technically coherent?
Rather than deleting sensitive content entirely, the implementation uses contextual placeholders that preserve grammatical structure and technical meaning. For example, company names become "申请人甲" (Applicant A), maintaining the semantic role of the entity in the sentence, while application numbers retain the CN prefix with masked digits to preserve format recognition. The de-identification rules specifically mandate that "the title's domain object must stay intact" while only surrounding identifying text undergoes redaction.
Can the de-identification rules be customized for specific patent portfolios?
Yes, the redact_text() function in skills/patent-oa/tools/redact.py accepts an extra_names parameter accepting a list of strings, allowing users to specify additional sensitive terms or internal codenames unique to their organization. Furthermore, the hash_app_nos boolean flag toggles between standardized placeholder masking (CNXXXXXXXXXX.X) and cryptographic hashing for application numbers, accommodating varying security requirements across different deployment scenarios.
Where is the de-identification logic integrated into the patent processing pipeline?
The redaction function operates at the ingestion boundary within skills/patent-oa/tools/ingest_case.py, ensuring that all patent disclosures undergo sanitization before entering the knowledge base or vector store. The test suite in skills/patent-oa/tests/test_oa_store.py validates that these transformations produce expected placeholder outputs, maintaining data quality standards throughout the storage workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →