How the Patent Plain Reader Extracts Claim Trees, Terminology, and Public Clues from PDFs

The Patent Plain Reader uses a three-stage pipeline: PyMuPDF-based text extraction in tools/oa/pdf_text.py, deterministic claim-tree parsing in tools/patent_reader/shared/common.py, and spaCy-powered terminology extraction plus regex-based public-clue detection in tools/patent_reader/analyze/lint_patent_note.py.

The handsomestWei/patent-disclosure-skill repository provides an open-source toolkit for transforming patent PDFs into structured, analyzable data. Understanding how the patent plain reader extracts claim trees, terminology, and public clues from PDFs reveals a robust pipeline designed for legal-tech workflows and prior-art research.

PDF Text Extraction with PyMuPDF

The pipeline begins in tools/oa/pdf_text.py, which wraps PyMuPDF (MuPDF) to convert patent documents into normalized Unicode strings.

The helper iterates page-by-page, extracting raw text while preserving layout hints such as font size changes and page numbers. These hints assist downstream tokenization. If a page fails text parsing, the module falls back to page-level OCR, ensuring no claim data is silently lost.

from tools.oa.pdf_text import pdf_to_text
from pathlib import Path

pdf_path = Path("sample_patent.pdf")
raw_text = pdf_to_text(pdf_path)  # Returns normalized Unicode with layout hints

This extraction layer handles the messy reality of patent PDFs—scanned images, mixed formats, and inconsistent encoding—before structured parsing begins.

Claim-Tree Construction and Validation

Once text is extracted, tools/patent_reader/shared/common.py transforms flat text into hierarchical claim structures.

The deterministic parser locates the "Claims" heading, then walks numbered lists (1., 2., 3.…). Using indentation patterns and punctuation, it constructs a tree structure where each claim node may reference parent claims (dependent claims) or stand as root claims (independent claims).

Two critical functions ensure data integrity:

  • normalize_claim_tree – Guarantees every claim has a parent field (None for roots) and correctly nests dependent claims
  • validate_claim_tree – Flags orphaned claim numbers, duplicate IDs, and missing metadata before downstream processing
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree

claim_tree = normalize_claim_tree(raw_text)   # Parses numbered claims into hierarchy

validation = validate_claim_tree(claim_tree)  # Sanity-checks structure

The validation step prevents malformed trees from reaching vault writers or analysis tools.

Terminology Extraction with spaCy

Domain-specific terminology is extracted by tools/patent_reader/analyze/lint_patent_note.py, which leverages spaCy's English model for noun-phrase detection.

The linter processes each claim node individually, pulling candidate terms that represent technical concepts, apparatus elements, or method steps. This terminological data supports downstream semantic analysis and cross-document comparison.

from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree

lint_result = lint_claim_tree(claim_tree)
print("Key terms:", lint_result["terms"])  # spaCy-extracted noun phrases

Public-Clue Detection via Regex Pattern Matching

The same linter scans claim text for citations to public patent documents—prior-art references that indicate relevant disclosures. The regex pattern:

r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b'

captures standard publication identifiers including:

  • US 1234567 A (US grants)
  • WO-2019/012345 (PCT applications)
  • EP-2 123 456 (European patents)

Matched identifiers populate a public_clues list that downstream tools render as searchable links.

print("Public clues found:", lint_result["public_clues"])

# Output: ['US 1234567 A', 'WO-2019/012345']

Complete Pipeline Integration

The end-to-end workflow ties extraction, parsing, and analysis together:

from pathlib import Path
from tools.oa.pdf_text import pdf_to_text
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree
from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree

pdf_path = Path("sample_patent.pdf")

# Stage 1: PDF → text

raw_text = pdf_to_text(pdf_path)

# Stage 2: Build and validate claim-tree

claim_tree = normalize_claim_tree(raw_text)
validation = validate_claim_tree(claim_tree)

# Stage 3: Extract terminology and public clues

lint_result = lint_claim_tree(claim_tree)

For Obsidian vault integration, tools/patent_reader/vault/write_patent_obsidian_note.py automates the full pipeline:

from tools.patent_reader.vault.write_patent_obsidian_note import write_note

write_note(
    pdf_path=Path("sample_patent.pdf"),
    workdir=Path("./my_vault"),
    public_clues=True,   # Embed extracted prior-art references

    terminology=True     # Embed extracted domain terms

)

Summary

The patent plain reader extracts claim trees, terminology, and public clues through three coordinated stages:

These components transform unstructured patent documents into machine-readable claim trees with enriched metadata for legal research workflows.

Frequently Asked Questions

What PDF library does the patent plain reader use?

The reader uses PyMuPDF (MuPDF) via tools/oa/pdf_text.py. This library provides fast page-by-page text extraction with access to font metadata and layout geometry. When text extraction fails, the module automatically falls back to OCR to preserve claim content.

How does the claim-tree parser handle dependent claims?

The deterministic parser in tools/patent_reader/shared/common.py identifies dependent claims through indentation patterns and textual parent references (e.g., "The apparatus of claim 1..."). It builds a tree where each node tracks its parent field, with root claims having parent: None. The normalize_claim_tree function ensures consistent nesting regardless of input formatting variations.

What types of public clues can the system detect?

The regex engine in lint_patent_note.py recognizes US, WO (PCT), and EP patent publication numbers in common formats. The pattern r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b' tolerates spacing variations and optional kind codes (A, B, etc.). These clues link to prior-art disclosures cited within claims or descriptions.

Can the pipeline integrate with note-taking systems beyond Obsidian?

The modular architecture supports custom vault writers. The write_patent_obsidian_note.py implementation demonstrates the interface: it consumes the standardized claim_tree JSON and lint_result dictionaries. Developers can implement alternative writers targeting Notion, Logseq, or internal databases using the same data structures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →