How to Use the Patent-Reader Skill for Patent Interpretation: A Complete Guide
The patent-reader skill transforms CNIPA publication numbers or local PDF files into structured Obsidian vaults containing hierarchical claim trees, figure assets, and interactive knowledge graphs through a self‑contained Python pipeline.
The patent-reader skill, housed in the handsomestWei/patent-disclosure-skill repository, automates patent interpretation by extracting technical content, building linked claim structures, and generating markdown notes without requiring external APIs. This tool is designed for IP researchers and patent attorneys who need to convert dense patent documents into navigable knowledge bases for prior‑art analysis and disclosure studies.
Architecture of the Patent-Reader Skill
The skill implements a four‑layer architecture that processes patent documents from raw input to structured output:
| Layer | Purpose | Primary Implementation |
|---|---|---|
| Input & Retrieval | Accepts publication numbers or PDF paths, downloads PDFs when necessary, and loads content into memory. | skills/patent-reader/tools/extract/fetch_patent_pdf.py |
| Content Extraction | Parses PDFs into plain text, claim hierarchies, figure assets, and term glossaries. | extract_patent_text.py, extract_patent_figures.py, figure_extract.py |
| Knowledge‑Graph Construction | Generates Obsidian vault markdown files and Canvas graphs linking claims, terms, and figures. | tools/vault/schema_vault.py, build_patent_canvas.py, write_patent_obsidian_note.py |
| Presentation & Export | Renders final markdown notes and optional DOCX/HTML reports. | tools/vault/obsidian.py, tools/vault/desc_paragraphs.py |
The Five‑Stage Interpretation Pipeline
When invoked, the skill executes a deterministic pipeline orchestrated through pure Python modules:
-
Environment Validation –
check_obsidian_env.pyverifies that thePATENT_READER_OBSIDIAN_VAULTenvironment variable points to a writable directory, falling back tooutputs/patent_reader/if unset. -
Fetch & Parse –
fetch_patent_pdf.pyretrieves the PDF from the CNIPA website when a publication number is provided, thenextract_patent_text.pyconverts the document into structured plain text. -
Claim Structuring –
validate_claim_tree.pyandbuild_claim_mermaid.pyparse dependent and independent claims into hierarchical trees and generate Mermaid diagrams for visualization. -
Note Generation –
write_patent_obsidian_note.pycreates markdown notes embedding the claim tree, term glossary, and figure references within the Obsidian vault. -
Canvas Visualization –
build_patent_canvas.pyproduces a.canvasfile that maps relationships between patents, individual claims, and public clues fetched byvalidate_public_clues.py.
Usage Methods for Patent Interpretation
Direct Python API Integration
Import the core functions to programmatically process patents within your own analysis scripts:
from skills.patent_reader.tools.extract.fetch_patent_pdf import fetch_patent_pdf
from skills.patent_reader.tools.vault.write_patent_obsidian_note import write_patent_note
# Interpret a CNIPA patent by publication number
pub_number = "CN1123456A"
pdf_path = fetch_patent_pdf(pub_number) # Downloads from CNIPA
write_patent_note(pdf_path, vault_path="~/Obsidian/PatentVault")
The fetch_patent_pdf function handles HTTP retrieval via cnipa_crawler.py, while write_patent_note triggers the full extraction pipeline and generates both the markdown note and the Canvas graph.
Instagit CLI Commands
Invoke the skill directly from the terminal using the Instagit ecosystem’s natural language interface:
# Interpret by publication number
skill invoke patent-reader "CN1123456A"
# Interpret a local PDF file
skill invoke patent-reader "/path/to/patent.pdf"
The CLI automatically opens the generated Obsidian note upon completion when the vault is properly configured. Alternatively, use the Mandarin command alias:
读专利 CN1123456A
Configuring the Obsidian Vault Path
Override the default output location by setting the environment variable before initialization:
import os
from skills.patent_reader.tools.vault.setup_obsidian_vault import setup_vault
# Redirect output to a custom directory
os.environ["PATENT_READER_OBSIDIAN_VAULT"] = "/tmp/patent_reader_output"
setup_vault() # Creates folder structure if absent
If PATENT_READER_OBSIDIAN_VAULT remains undefined, the skill automatically writes to outputs/patent_reader/ within the project root.
Core Implementation Files
The following modules contain the critical logic for patent interpretation:
fetch_patent_pdf.py– Downloads CNIPA PDFs given a publication number.extract_patent_text.py– Parses PDF content into raw text and structured claim trees.schema_vault.py– Defines the markdown schema and Canvas JSON structure for Obsidian integration.write_patent_obsidian_note.py– Orchestrates the pipeline and writes the final interpreted note.build_patent_canvas.py– Generates the interactive graph linking claims, figures, and external clues.test_patent_reader_pipeline.py– Provides end‑to‑end validation of the interpretation workflow.
Summary
- The patent-reader skill converts CNIPA numbers or PDFs into structured Obsidian vaults through a four‑layer Python pipeline.
- Execution requires no external databases; only optional internet access for fetching remote PDFs or public clues.
- Users can trigger interpretation via Python imports, Instagit CLI commands, or Mandarin language aliases.
- Output locations are configurable via the
PATENT_READER_OBSIDIAN_VAULTenvironment variable, defaulting tooutputs/patent_reader/.
Frequently Asked Questions
What input formats does the patent-reader skill support?
The skill accepts either a CNIPA publication number (e.g., CN1123456A) or a local file path pointing to a PDF document. When provided with a publication number, the system automatically downloads the corresponding PDF from the CNIPA public database before processing.
How does the skill handle CNIPA patent downloads?
The fetch_patent_pdf function in skills/patent-reader/tools/extract/fetch_patent_pdf.py crawls the CNIPA website using the cnipa_crawler.py helper, saves the PDF to a temporary location, and returns the local file path for subsequent extraction stages.
Can I use the patent-reader skill without Obsidian?
Yes. While the skill generates Obsidian‑compatible markdown and Canvas files by default, the underlying extraction logic in extract_patent_text.py and validate_claim_tree.py operates independently. You can consume the raw JSON or markdown outputs in any text editor or knowledge‑base system.
Where are the generated knowledge graphs stored?
The skill writes all output to the directory specified by the PATENT_READER_OBSIDIAN_VAULT environment variable. If this variable is unset, files are saved to outputs/patent_reader/ relative to the skill root, creating a portable vault structure containing the patent note, claim tree diagrams, and Canvas visualization files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →