# Patent-Reader Sub-Skill Toolchain Structure: A Complete Pipeline Guide

> Explore the patent-reader sub-skill toolchain structure. This guide details the four-layer pipeline for converting patent disclosures into Obsidian notes.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: architecture
- Published: 2026-09-05

---

**The patent-reader sub-skill uses a four-layer modular pipeline—environment setup, data acquisition, content processing, and output validation—to convert CNIPA patent publications into richly-formatted Obsidian notes.**

This article breaks down the complete toolchain structure for the **patent-reader** sub-skill in the [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill) repository. Each layer is implemented as a dedicated Python module under `skills/patent-reader/tools/`, designed for both manual CLI invocation and automated skill runtime execution.

## Environment & Setup Layer

Before any patent processing begins, the toolchain validates and prepares the Obsidian vault environment. Two scripts handle this bootstrap phase:

- **[`tools/vault/check_obsidian_env.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/check_obsidian_env.py)** — Verifies that environment variables point to a valid vault path and that required directories exist. Run with `--auto-accept` to create missing paths without interactive prompts.

- **[`tools/vault/setup_obsidian_vault.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/setup_obsidian_vault.py)** — Scaffolds the complete vault layout, including the [`patents.base.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/patents.base.yaml) configuration and the [`patent-reader.css`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/patent-reader.css) snippet located at [`assets/obsidian/patent-reader.css`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/assets/obsidian/patent-reader.css). This CSS file applies consistent styling to all generated patent notes via the `cssclasses: patent-reader` front-matter key.

```bash

# Initialize vault (run once)

python skills/patent-reader/tools/vault/setup_obsidian_vault.py \
    --vault /path/to/MyVault

```

## Data Acquisition Layer

This layer fetches raw patent artifacts from CNIPA (China National Intellectual Property Administration). The **patent-reader** toolchain supports three acquisition paths depending on patent type and available sources:

| Script | Purpose | Output Location |
|--------|---------|---------------|
| [`tools/extract/fetch_patent_pdf.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/extract/fetch_patent_pdf.py) | Downloads PDF by publication number | `outputs/patent_reader/<run>/patent.pdf` |
| [`tools/extract/fetch_design_views.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/extract/fetch_design_views.py) | Scrapes design-view PNG images | `outputs/patent_reader/<run>/design_views/` |
| [`tools/crawl/cnipa_epub_crawler.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/crawl/cnipa_epub_crawler.py) | Alternative EPUB crawler for CNIPA | Run-specific folder |

The fetch scripts populate a run-specific workspace under `outputs/patent_reader/<timestamp>/`, isolating each patent processing job.

```bash

# Fetch patent PDF

python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
    --pub CN119961396A \
    -o tmp/run

```

## Content Processing Layer

Raw artifacts are transformed into structured, analyzable data through six specialized tools. This is the most complex **patent-reader toolchain** layer, handling OCR, classification, and semantic analysis.

### Text and Figure Extraction

- **[`tools/extract/extract_patent_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/extract/extract_patent_text.py)** — Performs OCR on the PDF and converts content to markdown, preserving section hierarchy (claims, description, abstract).

- **[`tools/extract/extract_patent_figures.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/extract/extract_patent_figures.py)** — Identifies, extracts, and classifies figure images from the patent document.

### Patent Type Detection

- **[`tools/patent_type.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_type.py)** — Infers the patent category (invention/utility, utility model, or design) from the publication number format. Writes the classification to [`fetch_pdf_status.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/fetch_pdf_status.json) for downstream routing.

### Analysis and Visualization

- **[`tools/analyze/build_claim_mermaid.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/build_claim_mermaid.py)** — Generates a **Mermaid diagram** representing the claim dependency tree, saved as `.mmd` files for embedding in Obsidian.

- **[`tools/analyze/build_context_anchor.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/build_context_anchor.py)** — Creates context-anchor files that enable later cross-referencing between related patents.

- **[`tools/analyze/validate_claim_tree.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/validate_claim_tree.py)** & **[`tools/analyze/validate_public_clues.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/validate_public_clues.py)** — Sanity checks ensuring claim hierarchy integrity and coherence of public-clue JSON metadata.

```bash

# Full processing sequence

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i tmp/run/patent.pdf -o tmp/run

python skills/patent-reader/tools/patent_type.py --pub CN119961396A

python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
    --claim-tree tmp/run/claim_tree.json \
    --pub-number CN119961396A \
    -o tmp/run/claim_mermaid.mmd

```

## Output & Validation Layer

The final layer assembles, enriches, and validates the Obsidian-ready output. Four tools collaborate to produce publication-ready notes:

- **[`tools/vault/write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/write_patent_obsidian_note.py)** — Merges generated front-matter (`ipc`, `confidence_speculative`, `cssclasses: patent-reader`), schema JSON files ([`structure_schema.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/structure_schema.json), [`appearance_schema.json`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/appearance_schema.json)), and markdown body into the final note file.

- **[`tools/vault/build_patent_canvas.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/build_patent_canvas.py)** — Creates a **Canvas** aggregation page that visually groups related patent notes in a board layout.

- **[`tools/vault/link_patent_notes.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/vault/link_patent_notes.py)** — Auto-links notes by resolving `public-clue` references across the vault. Use `--dry-run` to preview connections before applying.

- **[`tools/analyze/lint_patent_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/analyze/lint_patent_note.py)** — Enforces schema compliance, verifying required front-matter keys and flagging malformed entries.

```bash

# Write final note

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --note-dir /path/to/MyVault/Research/Patents \
    --run-dir tmp/run

# Build Canvas and link related notes

python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/vault/link_patent_notes.py --dry-run
python skills/patent-reader/tools/vault/link_patent_notes.py

```

## Orchestration: The Complete Pipeline

The **patent-reader** toolchain executes as a linear pipeline, with each stage consuming outputs from the previous. The orchestrating prompt at [`prompts/patent_plain_reader.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/prompts/patent_plain_reader.md) drives this flow in skill runtime mode, but all components are independently runnable for debugging or custom workflows.

**Complete manual execution:**

```bash

# 1. Environment

python skills/patent-reader/tools/vault/check_obsidian_env.py --auto-accept

# 2. Acquisition

python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
    --pub CN119961396A -o tmp/run

# 3. Processing

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i tmp/run/patent.pdf -o tmp/run
python skills/patent-reader/tools/patent_type.py --pub CN119961396A
python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
    --claim-tree tmp/run/claim_tree.json \
    --pub-number CN119961396A -o tmp/run/claim_mermaid.mmd
python skills/patent-reader/tools/analyze/validate_claim_tree.py \
    --input tmp/run/claim_tree.json

# 4. Output

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --note-dir /path/to/MyVault/Research/Patents --run-dir tmp/run
python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/analyze/lint_patent_note.py \
    --input /path/to/MyVault/Research/Patents/CN119961396A.md

```

## Summary

- The **patent-reader toolchain structure** comprises four logical layers: environment setup, data acquisition, content processing, and output validation.

- All tools reside under `skills/patent-reader/tools/` with clear separation by function: `vault/` for Obsidian integration, `extract/` for data fetching and OCR, `analyze/` for claim processing and validation, and `crawl/` for alternative acquisition.

- Key entry points include [`setup_obsidian_vault.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/setup_obsidian_vault.py) for initialization, [`fetch_patent_pdf.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/fetch_patent_pdf.py) for acquisition, [`extract_patent_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/extract_patent_text.py) for content conversion, and [`write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/write_patent_obsidian_note.py) for final assembly.

- The pipeline generates **Mermaid claim diagrams**, **Canvas aggregation pages**, and **auto-linked note networks** tailored for patent research workflows.

## Frequently Asked Questions

### Where are the patent-reader tools located in the repository?

All sub-skill tools are under `skills/patent-reader/tools/`, organized by function: `vault/` for Obsidian environment management, `extract/` for PDF/text/figure processing, `analyze/` for claim validation and diagram generation, and `crawl/` for alternative EPUB acquisition. The orchestrating prompt lives at [`prompts/patent_plain_reader.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/prompts/patent_plain_reader.md).

### Can I run the patent-reader pipeline manually without the skill runtime?

Yes. Every tool is independently executable via CLI with `--help` for parameter documentation. The complete pipeline can be triggered step-by-step using the bash sequence shown in the orchestration section above, making it suitable for debugging or custom integration.

### What output formats does the patent-reader toolchain produce?

The pipeline generates three primary outputs: markdown notes with YAML front-matter for Obsidian, Mermaid diagram files (`.mmd`) for claim visualization, and Canvas JSON files for visual aggregation boards. All outputs include the `cssclasses: patent-reader` metadata for consistent styling.

### How does the toolchain handle different patent types?

The [`tools/patent_type.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_type.py) module analyzes publication number patterns to classify patents as invention/utility, utility model, or design. This classification routes processing logic—design patents trigger [`fetch_design_views.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/fetch_design_views.py) for image scraping, while utility patents emphasize claim tree analysis via [`build_claim_mermaid.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/build_claim_mermaid.py).