# How the patent-reader Sub-Skill Works: Automated Patent Processing for Obsidian Vaults

> Discover how the patent-reader sub-skill automates patent processing for Obsidian vaults. Learn about extraction, context anchoring, claim tree normalization, and vault writing stages.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-04

---

**The patent-reader sub-skill converts raw Chinese patent documents (PDF, Markdown, or plain-text) into structured, human-readable notes with visual Canvas navigation through a four-stage pipeline: extraction, context anchoring, claim tree normalization, and vault writing.**

The patent-reader sub-skill is a specialized component of the `handsomestWei/patent-disclosure-skill` repository designed to automate the ingestion and structuring of patent literature. It transforms unstructured legal documents into enriched Markdown files suitable for knowledge management systems, handling everything from PDF text extraction to Obsidian-compatible Canvas generation.

## Four-Stage Processing Pipeline

The patent-reader sub-skill orchestrates a **clean-extract → context-anchor → claim-tree → vault-write** workflow through four distinct stages wired together by command-line scripts and shared helper modules.

### Stage 1: Text Extraction and Parsing

The pipeline begins in [`tools/extract/extract_patent_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/extract/extract_patent_text.py), which handles the initial document ingestion and fragmentation.

When provided with a publication number, the system first downloads the full-text PDF using [`fetch_patent_pdf.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/fetch_patent_pdf.py) (internal helper). The `read_input()` function then opens the file, leveraging **PyMuPDF** (`fitz`) for PDF parsing. The extraction logic scans for critical patent sections:

- **`split_claims_block()`** isolates the claims section from the description
- **`extract_glossary()`** harvests terminology definitions
- **`extract_embodiments()`** captures specific implementation examples
- **`sections_from_text()`** assigns unique identifiers (e.g., `abstract`, `claim_1`, `desc_001`) to each fragment

The output produces two key artifacts: [`source_manifest.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/source_manifest.json) (containing publication numbers, claim counts, and IPC codes) and [`synthesis_bundle.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/synthesis_bundle.json) (aggregating sections, glossary candidates, and metadata).

### Stage 2: Context Anchor Generation

The [`tools/analyze/build_context_anchor.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/build_context_anchor.py) script consumes the extraction artifacts to generate [`context_anchor.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/context_anchor.json)—a technical anchor file that guides downstream processing.

Key functions in this stage include:

- **`resolve_domain()`** — Classifies the patent into technical domains using YAML rule sets
- **`resolve_ipc_hints()`** — Maps International Patent Classification codes to industry hints
- **`claim_keyword_tokens()`** — Derives search tokens from independent claims for semantic retrieval
- **`obsidian_navigation()`** — Generates web-search queries (e.g., "company product site:com") for external evidence gathering

This context anchor serves as the semantic backbone for subsequent chat-agent interactions and disclosure generation.

### Stage 3: Claim Tree Normalization

Raw patent claims require structural normalization to establish hierarchical relationships. The [`tools/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/shared/common.py) module provides the core logic for this transformation through **`normalize_claim_tree()`**.

The normalization process:

1. Identifies independent claims using **`guess_independent()`**
2. Maps parent-child relationships via **`parent_claim_numbers()`**
3. Attaches orphaned dependent claims to the nearest preceding independent claim
4. Detects and breaks circular dependencies
5. Validates structural integrity with **`validate_claim_tree()`**

This stage produces a clean, traversable claim hierarchy and optional Mermaid diagrams for visual representation.

### Stage 4: Obsidian Vault Integration

The final stage in [`tools/vault/write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/write_patent_obsidian_note.py) transforms lint-validated Markdown into fully-featured Obsidian vault entries.

The enrichment process includes:

- **`sanitize_user_facing_titles()`** — Standardizes heading formats for vault consistency
- **`enrich_note_frontmatter()`** — Injects metadata (domain, publication number, IPC hints) as YAML front-matter
- **`_inject_figure_embeds()`** — Copies figures to `images/` folders and inserts them into "### 附图" sections

- **`build_canvas()`** — Generates Obsidian Canvas JSON files for visual navigation
- **`ensure_canvas_nav()`** and **`upsert_index_entry()`** — Updates global and domain-specific Maps of Content (MOCs)
- **`ensure_source_pdf_nav()`** — Archives original PDFs in the vault's `source/` directory

When no Obsidian vault is detected, artifacts default to `outputs/patent_reader/` for local consumption.

## Key Implementation Details

### PDF Processing Architecture

The extraction layer relies on **PyMuPDF** for rendering Chinese patent PDFs into machine-readable text. The `read_input()` function handles format detection automatically, processing `.pdf`, `.md`, and `.txt` files through a unified interface.

### Claim Tree Logic

The claim normalization algorithm in [`common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/common.py) handles complex dependency scenarios:

- **Independent claims** are flagged and stripped of parent pointers
- **Dependent claims** with invalid parent references are "promoted" or reassigned
- **Cycle detection** prevents infinite loops in malformed claim sets

### Canvas Generation

The Obsidian Canvas construction (via [`obsidian.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/obsidian.py) and [`schema_vault.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/schema_vault.py)) creates interactive visual graphs linking patent sections, enabling non-linear navigation through complex disclosure documents.

## Command-Line Usage Examples

Execute the complete patent-reader workflow using these sequential commands:

```bash

# Stage 1: Extract structured text from patent PDF

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i CN123456789A.pdf \
    -o outputs/patent_reader/run1 \
    --pub-number CN123456789A

```

```bash

# Stage 2: Build context anchor with domain classification

python skills/patent-reader/tools/analyze/build_context_anchor.py \
    -w outputs/patent_reader/run1

```

```bash

# Stage 3: Validate note structure (optional linting step)

python skills/patent-reader/tools/analyze/lint_patent_note.py \
    --note my_note.md \
    --manifest outputs/patent_reader/run1/source_manifest.json \
    --claim-tree outputs/patent_reader/run1/claim_tree.json \
    --plan outputs/patent_reader/run1/note_plan.json \
    --context-anchor outputs/patent_reader/run1/context_anchor.json

```

```bash

# Stage 4: Write enriched note to Obsidian vault

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --content-file my_note.md \
    --manifest outputs/patent_reader/run1/source_manifest.json \
    --lint-json lint.json \
    --context-anchor outputs/patent_reader/run1/context_anchor.json \
    --bundle outputs/patent_reader/run1/synthesis_bundle.json \
    --workdir outputs/patent_reader/run1 \
    --output write_status.json

```

All commands assume execution from the repository root directory.

## Summary

- The **patent-reader sub-skill** automates the conversion of Chinese patent documents into structured knowledge base entries through four distinct stages.
- **Stage 1** ([`extract_patent_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/extract_patent_text.py)) handles PDF ingestion and section fragmentation using PyMuPDF.
- **Stage 2** ([`build_context_anchor.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/build_context_anchor.py)) generates semantic metadata including domain classification and IPC code mapping.
- **Stage 3** ([`common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/common.py)) normalizes claim hierarchies and resolves parent-child dependencies.
- **Stage 4** ([`write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/write_patent_obsidian_note.py)) produces Obsidian-compatible Markdown with embedded Canvas visualizations and navigation indexes.
- The system supports publication number-based PDF retrieval, glossary extraction, and automatic figure management.
- When vault integration is unavailable, the pipeline defaults to local file output in `outputs/patent_reader/`.

## Frequently Asked Questions

### What input formats does the patent-reader sub-skill support?

The patent-reader sub-skill accepts **PDF files**, **Markdown documents**, and **plain-text files**. For PDF processing, the system uses PyMuPDF (`fitz`) to extract text content. If only a publication number is provided (e.g., CN123456789A), the [`fetch_patent_pdf.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/fetch_patent_pdf.py) helper downloads the full-text PDF automatically before processing begins.

### How does the patent-reader handle claim dependencies?

The sub-skill employs `normalize_claim_tree()` in [`tools/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/shared/common.py) to resolve claim hierarchies. Independent claims are identified via `guess_independent()`, while `parent_claim_numbers()` maps dependent claims to their antecedents. The algorithm automatically attaches orphaned claims to the nearest valid independent claim and breaks circular references to ensure tree integrity.

### Can I use the patent-reader without an Obsidian vault?

Yes. When the script detects no Obsidian vault configuration, it writes all artifacts—including the enriched Markdown note, Canvas JSON, claim-tree side-files, and [`write_status.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/write_status.json)—to the local `outputs/patent_reader/` directory. The generated Markdown remains fully functional without vault-specific navigation features.

### What Python dependencies are required for the patent-reader?

The primary external dependency is **PyMuPDF** (`fitz`) for PDF text extraction. Additional requirements include standard libraries for JSON manipulation, YAML parsing for domain rules, and filesystem utilities. Refer to [`INSTALL.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/INSTALL.md) in the repository root for the complete dependency list and environment setup instructions.