# How the Patent Plain Reader Extracts Claim Trees, Terminology, and Public Clues from PDFs

> Discover how the Patent Plain Reader extracts claim trees, terminology, and public clues from PDFs. Learn about its three-stage pipeline including PyMuPDF, deterministic parsing, and spaCy analysis.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-02

---

**The Patent Plain Reader uses a three-stage pipeline: PyMuPDF-based text extraction in [`tools/oa/pdf_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/pdf_text.py), deterministic claim-tree parsing in [`tools/patent_reader/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/shared/common.py), and spaCy-powered terminology extraction plus regex-based public-clue detection in [`tools/patent_reader/analyze/lint_patent_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/analyze/lint_patent_note.py).**

The `handsomestWei/patent-disclosure-skill` repository provides an open-source toolkit for transforming patent PDFs into structured, analyzable data. Understanding how the patent plain reader extracts claim trees, terminology, and public clues from PDFs reveals a robust pipeline designed for legal-tech workflows and prior-art research.

## PDF Text Extraction with PyMuPDF

The pipeline begins in **[`tools/oa/pdf_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/pdf_text.py)**, which wraps **PyMuPDF** (MuPDF) to convert patent documents into normalized Unicode strings.

The helper iterates page-by-page, extracting raw text while preserving layout hints such as font size changes and page numbers. These hints assist downstream tokenization. If a page fails text parsing, the module falls back to **page-level OCR**, ensuring no claim data is silently lost.

```python
from tools.oa.pdf_text import pdf_to_text
from pathlib import Path

pdf_path = Path("sample_patent.pdf")
raw_text = pdf_to_text(pdf_path)  # Returns normalized Unicode with layout hints

```

This extraction layer handles the messy reality of patent PDFs—scanned images, mixed formats, and inconsistent encoding—before structured parsing begins.

## Claim-Tree Construction and Validation

Once text is extracted, **[`tools/patent_reader/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/shared/common.py)** transforms flat text into hierarchical claim structures.

The deterministic parser locates the "Claims" heading, then walks numbered lists (`1.`, `2.`, `3.`…). Using indentation patterns and punctuation, it constructs a **tree structure** where each claim node may reference parent claims (dependent claims) or stand as root claims (independent claims).

Two critical functions ensure data integrity:

- **`normalize_claim_tree`** – Guarantees every claim has a `parent` field (`None` for roots) and correctly nests dependent claims
- **`validate_claim_tree`** – Flags orphaned claim numbers, duplicate IDs, and missing metadata before downstream processing

```python
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree

claim_tree = normalize_claim_tree(raw_text)   # Parses numbered claims into hierarchy

validation = validate_claim_tree(claim_tree)  # Sanity-checks structure

```

The validation step prevents malformed trees from reaching vault writers or analysis tools.

## Terminology Extraction with spaCy

Domain-specific terminology is extracted by **[`tools/patent_reader/analyze/lint_patent_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/analyze/lint_patent_note.py)**, which leverages **spaCy's English model** for noun-phrase detection.

The linter processes each claim node individually, pulling candidate terms that represent technical concepts, apparatus elements, or method steps. This terminological data supports downstream semantic analysis and cross-document comparison.

```python
from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree

lint_result = lint_claim_tree(claim_tree)
print("Key terms:", lint_result["terms"])  # spaCy-extracted noun phrases

```

## Public-Clue Detection via Regex Pattern Matching

The same linter scans claim text for **citations to public patent documents**—prior-art references that indicate relevant disclosures. The regex pattern:

```python
r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b'

```

captures standard publication identifiers including:

- `US 1234567 A` (US grants)
- `WO-2019/012345` (PCT applications)
- `EP-2 123 456` (European patents)

Matched identifiers populate a **`public_clues`** list that downstream tools render as searchable links.

```python
print("Public clues found:", lint_result["public_clues"])

# Output: ['US 1234567 A', 'WO-2019/012345']

```

## Complete Pipeline Integration

The end-to-end workflow ties extraction, parsing, and analysis together:

```python
from pathlib import Path
from tools.oa.pdf_text import pdf_to_text
from tools.patent_reader.shared.common import normalize_claim_tree, validate_claim_tree
from tools.patent_reader.analyze.lint_patent_note import lint_claim_tree

pdf_path = Path("sample_patent.pdf")

# Stage 1: PDF → text

raw_text = pdf_to_text(pdf_path)

# Stage 2: Build and validate claim-tree

claim_tree = normalize_claim_tree(raw_text)
validation = validate_claim_tree(claim_tree)

# Stage 3: Extract terminology and public clues

lint_result = lint_claim_tree(claim_tree)

```

For Obsidian vault integration, **[`tools/patent_reader/vault/write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/vault/write_patent_obsidian_note.py)** automates the full pipeline:

```python
from tools.patent_reader.vault.write_patent_obsidian_note import write_note

write_note(
    pdf_path=Path("sample_patent.pdf"),
    workdir=Path("./my_vault"),
    public_clues=True,   # Embed extracted prior-art references

    terminology=True     # Embed extracted domain terms

)

```

## Summary

The patent plain reader extracts claim trees, terminology, and public clues through three coordinated stages:

- **PDF ingestion** – [`tools/oa/pdf_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/pdf_text.py) uses PyMuPDF with OCR fallback for reliable text extraction
- **Tree construction** – [`tools/patent_reader/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/shared/common.py) parses numbered claims into validated hierarchical structures
- **Semantic mining** – [`tools/patent_reader/analyze/lint_patent_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/analyze/lint_patent_note.py) applies spaCy for terminology and regex for public-clue detection

These components transform unstructured patent documents into machine-readable claim trees with enriched metadata for legal research workflows.

## Frequently Asked Questions

### What PDF library does the patent plain reader use?

The reader uses **PyMuPDF** (MuPDF) via [`tools/oa/pdf_text.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/pdf_text.py). This library provides fast page-by-page text extraction with access to font metadata and layout geometry. When text extraction fails, the module automatically falls back to OCR to preserve claim content.

### How does the claim-tree parser handle dependent claims?

The deterministic parser in [`tools/patent_reader/shared/common.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/patent_reader/shared/common.py) identifies dependent claims through indentation patterns and textual parent references (e.g., "The apparatus of claim 1..."). It builds a tree where each node tracks its `parent` field, with root claims having `parent: None`. The `normalize_claim_tree` function ensures consistent nesting regardless of input formatting variations.

### What types of public clues can the system detect?

The regex engine in [`lint_patent_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/lint_patent_note.py) recognizes **US**, **WO** (PCT), and **EP** patent publication numbers in common formats. The pattern `r'\b(?:US|WO|EP)[\s‑]?\d{1,7}[A-Z]?\b'` tolerates spacing variations and optional kind codes (A, B, etc.). These clues link to prior-art disclosures cited within claims or descriptions.

### Can the pipeline integrate with note-taking systems beyond Obsidian?

The modular architecture supports custom vault writers. The [`write_patent_obsidian_note.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/write_patent_obsidian_note.py) implementation demonstrates the interface: it consumes the standardized `claim_tree` JSON and `lint_result` dictionaries. Developers can implement alternative writers targeting Notion, Logseq, or internal databases using the same data structures.