Patent-Reader Sub-Skill Toolchain Structure: A Complete Pipeline Guide

The patent-reader sub-skill uses a four-layer modular pipeline—environment setup, data acquisition, content processing, and output validation—to convert CNIPA patent publications into richly-formatted Obsidian notes.

This article breaks down the complete toolchain structure for the patent-reader sub-skill in the handsomestWei/patent-disclosure-skill repository. Each layer is implemented as a dedicated Python module under skills/patent-reader/tools/, designed for both manual CLI invocation and automated skill runtime execution.

Environment & Setup Layer

Before any patent processing begins, the toolchain validates and prepares the Obsidian vault environment. Two scripts handle this bootstrap phase:


# Initialize vault (run once)

python skills/patent-reader/tools/vault/setup_obsidian_vault.py \
    --vault /path/to/MyVault

Data Acquisition Layer

This layer fetches raw patent artifacts from CNIPA (China National Intellectual Property Administration). The patent-reader toolchain supports three acquisition paths depending on patent type and available sources:

Script Purpose Output Location
tools/extract/fetch_patent_pdf.py Downloads PDF by publication number outputs/patent_reader/<run>/patent.pdf
tools/extract/fetch_design_views.py Scrapes design-view PNG images outputs/patent_reader/<run>/design_views/
tools/crawl/cnipa_epub_crawler.py Alternative EPUB crawler for CNIPA Run-specific folder

The fetch scripts populate a run-specific workspace under outputs/patent_reader/<timestamp>/, isolating each patent processing job.


# Fetch patent PDF

python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
    --pub CN119961396A \
    -o tmp/run

Content Processing Layer

Raw artifacts are transformed into structured, analyzable data through six specialized tools. This is the most complex patent-reader toolchain layer, handling OCR, classification, and semantic analysis.

Text and Figure Extraction

Patent Type Detection

  • tools/patent_type.py — Infers the patent category (invention/utility, utility model, or design) from the publication number format. Writes the classification to fetch_pdf_status.json for downstream routing.

Analysis and Visualization


# Full processing sequence

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i tmp/run/patent.pdf -o tmp/run

python skills/patent-reader/tools/patent_type.py --pub CN119961396A

python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
    --claim-tree tmp/run/claim_tree.json \
    --pub-number CN119961396A \
    -o tmp/run/claim_mermaid.mmd

Output & Validation Layer

The final layer assembles, enriches, and validates the Obsidian-ready output. Four tools collaborate to produce publication-ready notes:


# Write final note

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --note-dir /path/to/MyVault/Research/Patents \
    --run-dir tmp/run

# Build Canvas and link related notes

python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/vault/link_patent_notes.py --dry-run
python skills/patent-reader/tools/vault/link_patent_notes.py

Orchestration: The Complete Pipeline

The patent-reader toolchain executes as a linear pipeline, with each stage consuming outputs from the previous. The orchestrating prompt at prompts/patent_plain_reader.md drives this flow in skill runtime mode, but all components are independently runnable for debugging or custom workflows.

Complete manual execution:


# 1. Environment

python skills/patent-reader/tools/vault/check_obsidian_env.py --auto-accept

# 2. Acquisition

python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
    --pub CN119961396A -o tmp/run

# 3. Processing

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i tmp/run/patent.pdf -o tmp/run
python skills/patent-reader/tools/patent_type.py --pub CN119961396A
python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
    --claim-tree tmp/run/claim_tree.json \
    --pub-number CN119961396A -o tmp/run/claim_mermaid.mmd
python skills/patent-reader/tools/analyze/validate_claim_tree.py \
    --input tmp/run/claim_tree.json

# 4. Output

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --note-dir /path/to/MyVault/Research/Patents --run-dir tmp/run
python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/analyze/lint_patent_note.py \
    --input /path/to/MyVault/Research/Patents/CN119961396A.md

Summary

  • The patent-reader toolchain structure comprises four logical layers: environment setup, data acquisition, content processing, and output validation.

  • All tools reside under skills/patent-reader/tools/ with clear separation by function: vault/ for Obsidian integration, extract/ for data fetching and OCR, analyze/ for claim processing and validation, and crawl/ for alternative acquisition.

  • Key entry points include setup_obsidian_vault.py for initialization, fetch_patent_pdf.py for acquisition, extract_patent_text.py for content conversion, and write_patent_obsidian_note.py for final assembly.

  • The pipeline generates Mermaid claim diagrams, Canvas aggregation pages, and auto-linked note networks tailored for patent research workflows.

Frequently Asked Questions

Where are the patent-reader tools located in the repository?

All sub-skill tools are under skills/patent-reader/tools/, organized by function: vault/ for Obsidian environment management, extract/ for PDF/text/figure processing, analyze/ for claim validation and diagram generation, and crawl/ for alternative EPUB acquisition. The orchestrating prompt lives at prompts/patent_plain_reader.md.

Can I run the patent-reader pipeline manually without the skill runtime?

Yes. Every tool is independently executable via CLI with --help for parameter documentation. The complete pipeline can be triggered step-by-step using the bash sequence shown in the orchestration section above, making it suitable for debugging or custom integration.

What output formats does the patent-reader toolchain produce?

The pipeline generates three primary outputs: markdown notes with YAML front-matter for Obsidian, Mermaid diagram files (.mmd) for claim visualization, and Canvas JSON files for visual aggregation boards. All outputs include the cssclasses: patent-reader metadata for consistent styling.

How does the toolchain handle different patent types?

The tools/patent_type.py module analyzes publication number patterns to classify patents as invention/utility, utility model, or design. This classification routes processing logic—design patents trigger fetch_design_views.py for image scraping, while utility patents emphasize claim tree analysis via build_claim_mermaid.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →