How the Patent Plain Reader Extracts Claim Trees, Terminology, and Public Clues from PDFs
The Patent Plain Reader uses a three-stage pipeline: PyMuPDF-based text extraction in tools/oa/pdf_text.py, deterministic claim-tree parsing in tools/patent_reader/shared/common.py, and spaCy-powered terminology extraction plus regex-based public-clue detection in tools/patent_reader/analyze/lint_patent_note.py.
The handsomestWei/patent-disclosure-skill repository provides an open-source toolkit for transforming patent PDFs into structured, analyzable data. Understanding how the patent plain reader extracts claim trees, terminology, and public clues from PDFs reveals a robust pipeline designed for legal-tech workflows and prior-art research.
PDF Text Extraction with PyMuPDF
The pipeline begins in tools/oa/pdf_text.py, which wraps PyMuPDF (MuPDF) to convert patent documents into normalized Unicode strings.
The helper iterates page-by-page, extracting raw text while preserving layout hints such as font size changes and page numbers. These hints assist downstream tokenization. If a page fails text parsing, the module falls back to page-level OCR, ensuring no claim data is silently lost.
from tools.oa.pdf_text import pdf_to_text
from pathlib import Path
pdf_path = Path("sample_patent.pdf")
raw_text = pdf_to_text(pdf_path) # Returns normalized Unicode with layout hints
This extraction layer handles the messy reality of patent PDFs—scanned images, mixed formats, and inconsistent encoding—before structured parsing begins.
Claim-Tree Construction and Validation
Once text is extracted, tools/patent_reader/shared/common.py transforms flat text into hierarchical claim structures.
The deterministic parser locates the "Claims" heading, then walks numbered lists (1., 2., 3.…). Using indentation patterns and punctuation, it constructs a tree structure where each claim node may reference parent claims (dependent claims) or stand as root claims (independent claims).
Two critical functions ensure data integrity:
normalize_claim_tree– Guarantees every claim has aparentfield (Nonefor roots) and correctly nests dependent claimsvalidate_claim_tree– Flags orphaned claim numbers, duplicate IDs, and missing metadata before downstream processing
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree
claim_tree = normalize_claim_tree(raw_text) # Parses numbered claims into hierarchy
validation = validate_claim_tree(claim_tree) # Sanity-checks structure
The validation step prevents malformed trees from reaching vault writers or analysis tools.
Terminology Extraction with spaCy
Domain-specific terminology is extracted by tools/patent_reader/analyze/lint_patent_note.py, which leverages spaCy's English model for noun-phrase detection.
The linter processes each claim node individually, pulling candidate terms that represent technical concepts, apparatus elements, or method steps. This terminological data supports downstream semantic analysis and cross-document comparison.
from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree
lint_result = lint_claim_tree(claim_tree)
print("Key terms:", lint_result["terms"]) # spaCy-extracted noun phrases
Public-Clue Detection via Regex Pattern Matching
The same linter scans claim text for citations to public patent documents—prior-art references that indicate relevant disclosures. The regex pattern:
r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b'
captures standard publication identifiers including:
US 1234567 A(US grants)WO-2019/012345(PCT applications)EP-2 123 456(European patents)
Matched identifiers populate a public_clues list that downstream tools render as searchable links.
print("Public clues found:", lint_result["public_clues"])
# Output: ['US 1234567 A', 'WO-2019/012345']
Complete Pipeline Integration
The end-to-end workflow ties extraction, parsing, and analysis together:
from pathlib import Path
from tools.oa.pdf_text import pdf_to_text
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree
from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree
pdf_path = Path("sample_patent.pdf")
# Stage 1: PDF → text
raw_text = pdf_to_text(pdf_path)
# Stage 2: Build and validate claim-tree
claim_tree = normalize_claim_tree(raw_text)
validation = validate_claim_tree(claim_tree)
# Stage 3: Extract terminology and public clues
lint_result = lint_claim_tree(claim_tree)
For Obsidian vault integration, tools/patent_reader/vault/write_patent_obsidian_note.py automates the full pipeline:
from tools.patent_reader.vault.write_patent_obsidian_note import write_note
write_note(
pdf_path=Path("sample_patent.pdf"),
workdir=Path("./my_vault"),
public_clues=True, # Embed extracted prior-art references
terminology=True # Embed extracted domain terms
)
Summary
The patent plain reader extracts claim trees, terminology, and public clues through three coordinated stages:
- PDF ingestion –
tools/oa/pdf_text.pyuses PyMuPDF with OCR fallback for reliable text extraction - Tree construction –
tools/patent_reader/shared/common.pyparses numbered claims into validated hierarchical structures - Semantic mining –
tools/patent_reader/analyze/lint_patent_note.pyapplies spaCy for terminology and regex for public-clue detection
These components transform unstructured patent documents into machine-readable claim trees with enriched metadata for legal research workflows.
Frequently Asked Questions
What PDF library does the patent plain reader use?
The reader uses PyMuPDF (MuPDF) via tools/oa/pdf_text.py. This library provides fast page-by-page text extraction with access to font metadata and layout geometry. When text extraction fails, the module automatically falls back to OCR to preserve claim content.
How does the claim-tree parser handle dependent claims?
The deterministic parser in tools/patent_reader/shared/common.py identifies dependent claims through indentation patterns and textual parent references (e.g., "The apparatus of claim 1..."). It builds a tree where each node tracks its parent field, with root claims having parent: None. The normalize_claim_tree function ensures consistent nesting regardless of input formatting variations.
What types of public clues can the system detect?
The regex engine in lint_patent_note.py recognizes US, WO (PCT), and EP patent publication numbers in common formats. The pattern r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b' tolerates spacing variations and optional kind codes (A, B, etc.). These clues link to prior-art disclosures cited within claims or descriptions.
Can the pipeline integrate with note-taking systems beyond Obsidian?
The modular architecture supports custom vault writers. The write_patent_obsidian_note.py implementation demonstrates the interface: it consumes the standardized claim_tree JSON and lint_result dictionaries. Developers can implement alternative writers targeting Notion, Logseq, or internal databases using the same data structures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →