How the patent-reader Sub-Skill Works: Automated Patent Processing for Obsidian Vaults

The patent-reader sub-skill converts raw Chinese patent documents (PDF, Markdown, or plain-text) into structured, human-readable notes with visual Canvas navigation through a four-stage pipeline: extraction, context anchoring, claim tree normalization, and vault writing.

The patent-reader sub-skill is a specialized component of the handsomestWei/patent-disclosure-skill repository designed to automate the ingestion and structuring of patent literature. It transforms unstructured legal documents into enriched Markdown files suitable for knowledge management systems, handling everything from PDF text extraction to Obsidian-compatible Canvas generation.

Four-Stage Processing Pipeline

The patent-reader sub-skill orchestrates a clean-extract → context-anchor → claim-tree → vault-write workflow through four distinct stages wired together by command-line scripts and shared helper modules.

Stage 1: Text Extraction and Parsing

The pipeline begins in tools/extract/extract_patent_text.py, which handles the initial document ingestion and fragmentation.

When provided with a publication number, the system first downloads the full-text PDF using fetch_patent_pdf.py (internal helper). The read_input() function then opens the file, leveraging PyMuPDF (fitz) for PDF parsing. The extraction logic scans for critical patent sections:

  • split_claims_block() isolates the claims section from the description
  • extract_glossary() harvests terminology definitions
  • extract_embodiments() captures specific implementation examples
  • sections_from_text() assigns unique identifiers (e.g., abstract, claim_1, desc_001) to each fragment

The output produces two key artifacts: source_manifest.json (containing publication numbers, claim counts, and IPC codes) and synthesis_bundle.json (aggregating sections, glossary candidates, and metadata).

Stage 2: Context Anchor Generation

The tools/analyze/build_context_anchor.py script consumes the extraction artifacts to generate context_anchor.json—a technical anchor file that guides downstream processing.

Key functions in this stage include:

  • resolve_domain() — Classifies the patent into technical domains using YAML rule sets
  • resolve_ipc_hints() — Maps International Patent Classification codes to industry hints
  • claim_keyword_tokens() — Derives search tokens from independent claims for semantic retrieval
  • obsidian_navigation() — Generates web-search queries (e.g., "company product site:com") for external evidence gathering

This context anchor serves as the semantic backbone for subsequent chat-agent interactions and disclosure generation.

Stage 3: Claim Tree Normalization

Raw patent claims require structural normalization to establish hierarchical relationships. The tools/shared/common.py module provides the core logic for this transformation through normalize_claim_tree().

The normalization process:

  1. Identifies independent claims using guess_independent()
  2. Maps parent-child relationships via parent_claim_numbers()
  3. Attaches orphaned dependent claims to the nearest preceding independent claim
  4. Detects and breaks circular dependencies
  5. Validates structural integrity with validate_claim_tree()

This stage produces a clean, traversable claim hierarchy and optional Mermaid diagrams for visual representation.

Stage 4: Obsidian Vault Integration

The final stage in tools/vault/write_patent_obsidian_note.py transforms lint-validated Markdown into fully-featured Obsidian vault entries.

The enrichment process includes:

  • sanitize_user_facing_titles() — Standardizes heading formats for vault consistency

  • enrich_note_frontmatter() — Injects metadata (domain, publication number, IPC hints) as YAML front-matter

  • _inject_figure_embeds() — Copies figures to images/ folders and inserts them into "### 附图" sections

  • build_canvas() — Generates Obsidian Canvas JSON files for visual navigation

  • ensure_canvas_nav() and upsert_index_entry() — Updates global and domain-specific Maps of Content (MOCs)

  • ensure_source_pdf_nav() — Archives original PDFs in the vault's source/ directory

When no Obsidian vault is detected, artifacts default to outputs/patent_reader/ for local consumption.

Key Implementation Details

PDF Processing Architecture

The extraction layer relies on PyMuPDF for rendering Chinese patent PDFs into machine-readable text. The read_input() function handles format detection automatically, processing .pdf, .md, and .txt files through a unified interface.

Claim Tree Logic

The claim normalization algorithm in common.py handles complex dependency scenarios:

  • Independent claims are flagged and stripped of parent pointers
  • Dependent claims with invalid parent references are "promoted" or reassigned
  • Cycle detection prevents infinite loops in malformed claim sets

Canvas Generation

The Obsidian Canvas construction (via obsidian.py and schema_vault.py) creates interactive visual graphs linking patent sections, enabling non-linear navigation through complex disclosure documents.

Command-Line Usage Examples

Execute the complete patent-reader workflow using these sequential commands:


# Stage 1: Extract structured text from patent PDF

python skills/patent-reader/tools/extract/extract_patent_text.py \
    -i CN123456789A.pdf \
    -o outputs/patent_reader/run1 \
    --pub-number CN123456789A

# Stage 2: Build context anchor with domain classification

python skills/patent-reader/tools/analyze/build_context_anchor.py \
    -w outputs/patent_reader/run1

# Stage 3: Validate note structure (optional linting step)

python skills/patent-reader/tools/analyze/lint_patent_note.py \
    --note my_note.md \
    --manifest outputs/patent_reader/run1/source_manifest.json \
    --claim-tree outputs/patent_reader/run1/claim_tree.json \
    --plan outputs/patent_reader/run1/note_plan.json \
    --context-anchor outputs/patent_reader/run1/context_anchor.json

# Stage 4: Write enriched note to Obsidian vault

python skills/patent-reader/tools/vault/write_patent_obsidian_note.py \
    --content-file my_note.md \
    --manifest outputs/patent_reader/run1/source_manifest.json \
    --lint-json lint.json \
    --context-anchor outputs/patent_reader/run1/context_anchor.json \
    --bundle outputs/patent_reader/run1/synthesis_bundle.json \
    --workdir outputs/patent_reader/run1 \
    --output write_status.json

All commands assume execution from the repository root directory.

Summary

  • The patent-reader sub-skill automates the conversion of Chinese patent documents into structured knowledge base entries through four distinct stages.
  • Stage 1 (extract_patent_text.py) handles PDF ingestion and section fragmentation using PyMuPDF.
  • Stage 2 (build_context_anchor.py) generates semantic metadata including domain classification and IPC code mapping.
  • Stage 3 (common.py) normalizes claim hierarchies and resolves parent-child dependencies.
  • Stage 4 (write_patent_obsidian_note.py) produces Obsidian-compatible Markdown with embedded Canvas visualizations and navigation indexes.
  • The system supports publication number-based PDF retrieval, glossary extraction, and automatic figure management.
  • When vault integration is unavailable, the pipeline defaults to local file output in outputs/patent_reader/.

Frequently Asked Questions

What input formats does the patent-reader sub-skill support?

The patent-reader sub-skill accepts PDF files, Markdown documents, and plain-text files. For PDF processing, the system uses PyMuPDF (fitz) to extract text content. If only a publication number is provided (e.g., CN123456789A), the fetch_patent_pdf.py helper downloads the full-text PDF automatically before processing begins.

How does the patent-reader handle claim dependencies?

The sub-skill employs normalize_claim_tree() in tools/shared/common.py to resolve claim hierarchies. Independent claims are identified via guess_independent(), while parent_claim_numbers() maps dependent claims to their antecedents. The algorithm automatically attaches orphaned claims to the nearest valid independent claim and breaks circular references to ensure tree integrity.

Can I use the patent-reader without an Obsidian vault?

Yes. When the script detects no Obsidian vault configuration, it writes all artifacts—including the enriched Markdown note, Canvas JSON, claim-tree side-files, and write_status.json—to the local outputs/patent_reader/ directory. The generated Markdown remains fully functional without vault-specific navigation features.

What Python dependencies are required for the patent-reader?

The primary external dependency is PyMuPDF (fitz) for PDF text extraction. Additional requirements include standard libraries for JSON manipulation, YAML parsing for domain rules, and filesystem utilities. Refer to INSTALL.md in the repository root for the complete dependency list and environment setup instructions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →