Patent-Reader Sub-Skill Toolchain Structure: A Complete Pipeline Guide
The patent-reader sub-skill uses a four-layer modular pipeline—environment setup, data acquisition, content processing, and output validation—to convert CNIPA patent publications into richly-formatted Obsidian notes.
This article breaks down the complete toolchain structure for the patent-reader sub-skill in the handsomestWei/patent-disclosure-skill repository. Each layer is implemented as a dedicated Python module under skills/patent-reader/tools/, designed for both manual CLI invocation and automated skill runtime execution.
Environment & Setup Layer
Before any patent processing begins, the toolchain validates and prepares the Obsidian vault environment. Two scripts handle this bootstrap phase:
-
tools/vault/check_obsidian_env.py— Verifies that environment variables point to a valid vault path and that required directories exist. Run with--auto-acceptto create missing paths without interactive prompts. -
tools/vault/setup_obsidian_vault.py— Scaffolds the complete vault layout, including thepatents.base.yamlconfiguration and thepatent-reader.csssnippet located atassets/obsidian/patent-reader.css. This CSS file applies consistent styling to all generated patent notes via thecssclasses: patent-readerfront-matter key.
# Initialize vault (run once)
python skills/patent-reader/tools/vault/setup_obsidian_vault.py \
--vault /path/to/MyVault
Data Acquisition Layer
This layer fetches raw patent artifacts from CNIPA (China National Intellectual Property Administration). The patent-reader toolchain supports three acquisition paths depending on patent type and available sources:
| Script | Purpose | Output Location |
|---|---|---|
tools/extract/fetch_patent_pdf.py |
Downloads PDF by publication number | outputs/patent_reader/<run>/patent.pdf |
tools/extract/fetch_design_views.py |
Scrapes design-view PNG images | outputs/patent_reader/<run>/design_views/ |
tools/crawl/cnipa_epub_crawler.py |
Alternative EPUB crawler for CNIPA | Run-specific folder |
The fetch scripts populate a run-specific workspace under outputs/patent_reader/<timestamp>/, isolating each patent processing job.
# Fetch patent PDF
python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
--pub CN119961396A \
-o tmp/run
Content Processing Layer
Raw artifacts are transformed into structured, analyzable data through six specialized tools. This is the most complex patent-reader toolchain layer, handling OCR, classification, and semantic analysis.
Text and Figure Extraction
-
tools/extract/extract_patent_text.py— Performs OCR on the PDF and converts content to markdown, preserving section hierarchy (claims, description, abstract). -
tools/extract/extract_patent_figures.py— Identifies, extracts, and classifies figure images from the patent document.
Patent Type Detection
tools/patent_type.py— Infers the patent category (invention/utility, utility model, or design) from the publication number format. Writes the classification tofetch_pdf_status.jsonfor downstream routing.
Analysis and Visualization
-
tools/analyze/build_claim_mermaid.py— Generates a Mermaid diagram representing the claim dependency tree, saved as.mmdfiles for embedding in Obsidian. -
tools/analyze/build_context_anchor.py— Creates context-anchor files that enable later cross-referencing between related patents. -
tools/analyze/validate_claim_tree.py&tools/analyze/validate_public_clues.py— Sanity checks ensuring claim hierarchy integrity and coherence of public-clue JSON metadata.
# Full processing sequence
python skills/patent-reader/tools/extract/extract_patent_text.py \
-i tmp/run/patent.pdf -o tmp/run
python skills/patent-reader/tools/patent_type.py --pub CN119961396A
python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
--claim-tree tmp/run/claim_tree.json \
--pub-number CN119961396A \
-o tmp/run/claim_mermaid.mmd
Output & Validation Layer
The final layer assembles, enriches, and validates the Obsidian-ready output. Four tools collaborate to produce publication-ready notes:
-
tools/vault/write_patent_obsidian_note.py— Merges generated front-matter (ipc,confidence_speculative,cssclasses: patent-reader), schema JSON files (structure_schema.json,appearance_schema.json), and markdown body into the final note file. -
tools/vault/build_patent_canvas.py— Creates a Canvas aggregation page that visually groups related patent notes in a board layout. -
tools/vault/link_patent_notes.py— Auto-links notes by resolvingpublic-cluereferences across the vault. Use--dry-runto preview connections before applying. -
tools/analyze/lint_patent_note.py— Enforces schema compliance, verifying required front-matter keys and flagging malformed entries.
# Write final note
python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
--note-dir /path/to/MyVault/Research/Patents \
--run-dir tmp/run
# Build Canvas and link related notes
python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/vault/link_patent_notes.py --dry-run
python skills/patent-reader/tools/vault/link_patent_notes.py
Orchestration: The Complete Pipeline
The patent-reader toolchain executes as a linear pipeline, with each stage consuming outputs from the previous. The orchestrating prompt at prompts/patent_plain_reader.md drives this flow in skill runtime mode, but all components are independently runnable for debugging or custom workflows.
Complete manual execution:
# 1. Environment
python skills/patent-reader/tools/vault/check_obsidian_env.py --auto-accept
# 2. Acquisition
python skills/patent-reader/tools/extract/fetch_patent_pdf.py \
--pub CN119961396A -o tmp/run
# 3. Processing
python skills/patent-reader/tools/extract/extract_patent_text.py \
-i tmp/run/patent.pdf -o tmp/run
python skills/patent-reader/tools/patent_type.py --pub CN119961396A
python skills/patent-reader/tools/analyze/build_claim_mermaid.py \
--claim-tree tmp/run/claim_tree.json \
--pub-number CN119961396A -o tmp/run/claim_mermaid.mmd
python skills/patent-reader/tools/analyze/validate_claim_tree.py \
--input tmp/run/claim_tree.json
# 4. Output
python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
--note-dir /path/to/MyVault/Research/Patents --run-dir tmp/run
python skills/patent-reader/tools/vault/build_patent_canvas.py -w tmp/run
python skills/patent-reader/tools/analyze/lint_patent_note.py \
--input /path/to/MyVault/Research/Patents/CN119961396A.md
Summary
-
The patent-reader toolchain structure comprises four logical layers: environment setup, data acquisition, content processing, and output validation.
-
All tools reside under
skills/patent-reader/tools/with clear separation by function:vault/for Obsidian integration,extract/for data fetching and OCR,analyze/for claim processing and validation, andcrawl/for alternative acquisition. -
Key entry points include
setup_obsidian_vault.pyfor initialization,fetch_patent_pdf.pyfor acquisition,extract_patent_text.pyfor content conversion, andwrite_patent_obsidian_note.pyfor final assembly. -
The pipeline generates Mermaid claim diagrams, Canvas aggregation pages, and auto-linked note networks tailored for patent research workflows.
Frequently Asked Questions
Where are the patent-reader tools located in the repository?
All sub-skill tools are under skills/patent-reader/tools/, organized by function: vault/ for Obsidian environment management, extract/ for PDF/text/figure processing, analyze/ for claim validation and diagram generation, and crawl/ for alternative EPUB acquisition. The orchestrating prompt lives at prompts/patent_plain_reader.md.
Can I run the patent-reader pipeline manually without the skill runtime?
Yes. Every tool is independently executable via CLI with --help for parameter documentation. The complete pipeline can be triggered step-by-step using the bash sequence shown in the orchestration section above, making it suitable for debugging or custom integration.
What output formats does the patent-reader toolchain produce?
The pipeline generates three primary outputs: markdown notes with YAML front-matter for Obsidian, Mermaid diagram files (.mmd) for claim visualization, and Canvas JSON files for visual aggregation boards. All outputs include the cssclasses: patent-reader metadata for consistent styling.
How does the toolchain handle different patent types?
The tools/patent_type.py module analyzes publication number patterns to classify patents as invention/utility, utility model, or design. This classification routes processing logic—design patents trigger fetch_design_views.py for image scraping, while utility patents emphasize claim tree analysis via build_claim_mermaid.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →