# Core Python Dependencies for patent-disclosure-skill: Complete Technical Breakdown

> Discover core Python dependencies for patent-disclosure-skill including python-docx, latex2mathml, PyYAML, playwright, mammoth, and python-pptx. Get a full technical breakdown for your project.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: technical-breakdown
- Published: 2026-09-05

---

**The core Python dependencies for patent-disclosure-skill are `python-docx`, `latex2mathml`, `PyYAML`, `playwright`, `mammoth`, and `python-pptx`, each providing essential document generation, math rendering, configuration management, web automation, and presentation capabilities.**

This article examines the six foundational packages that power the **patent-disclosure-skill** repository by handsomestWei. These dependencies enable patent search automation, disclosure document generation, and Office Action analysis through a carefully curated stack focused on document processing and web interaction.

## Dependency Overview from requirements.txt

The project's [`requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt) at the repository root defines the minimum versions and purposes of each core package. Below is the complete breakdown with implementation details.

### python-docx (≥1.1.0): Word Document Generation

The `python-docx` library handles all Microsoft Word (.docx) document creation and parsing. In [`skills/patent-disclosure/tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/md_to_docx.py), this dependency generates patent disclosure documents and OA opinion reports with precise formatting control.

```python
from docx import Document
from docx.shared import Inches, Pt

def create_disclosure(title: str, content: str) -> Document:
    doc = Document()
    heading = doc.add_heading(title, level=0)
    heading.alignment = 1  # Center alignment

    
    paragraph = doc.add_paragraph(content)
    paragraph_format = paragraph.paragraph_format
    paragraph_format.line_spacing = 1.15
    
    return doc

```

Source reference: [requirements.txt L3–L4](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L3)

### latex2mathml (≥3.77.0): Mathematical Formula Rendering

Patent documents frequently contain complex mathematical formulas. The `latex2mathml` package converts LaTeX expressions to MathML, which then transforms to Office Math Markup Language (OMML) for native Word rendering.

```python
from latex2mathml import latex2mathml
from docx import Document

def insert_equation(doc: Document, latex_expr: str):
    # Convert LaTeX → MathML string

    mathml_string = latex2mathml(latex_expr)
    
    # Add to document (python-docx handles OMML conversion)

    paragraph = doc.add_paragraph()
    run = paragraph.add_run()
    run._element.append(mathml_to_omml_element(mathml_string))
    
    return doc

```

This integration appears throughout [`skills/patent-disclosure/tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/md_to_docx.py) for rendering patent claim formulas and technical equations.

Source reference: [requirements.txt L5–L6](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L5)

### PyYAML (≥6.0): Configuration and Schema Management

Configuration loading uses `PyYAML` throughout the skill modules. The repository employs YAML files for paradigm definitions, tool configurations, and schema specifications.

```python
import yaml
from pathlib import Path

def load_config(config_path: Path = Path("config.yaml")):
    with open(config_path, "r", encoding="utf-8") as f:
        config = yaml.safe_load(f)
    
    # Validate required sections

    required_keys = ["disclosure_settings", "search_parameters"]
    for key in required_keys:
        if key not in config:
            raise ValueError(f"Missing required config key: {key}")
    
    return config

def load_paradigms(paradigm_path: Path = Path("paradigms.yaml")):
    with open(paradigm_path, "r", encoding="utf-8") as f:
        return yaml.safe_load(f)

```

The [`skills/patent-disclosure/tools/config.yaml`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/config.yaml) file demonstrates this pattern, storing disclosure generation parameters and search API endpoints.

Source reference: [requirements.txt L7–L8](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L7)

### playwright (≥1.40.0): Headless Browser Automation

Web-based patent searches leverage `playwright` for automated browser control. The [`skills/patent-search/tools/browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/browser.py) module implements CNIPA (China National Intellectual Property Administration) searches with this framework.

```python
from playwright.sync_api import sync_playwright, Page, Browser
from typing import List, Dict
import time

def cnipa_bibliographic_search(
    application_number: str,
    headless: bool = True
) -> Dict[str, str]:
    results = {}
    
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=headless)
        context = browser.new_context(
            user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.0"
        )
        page = context.new_page()
        
        # Navigate to CNIPA search portal

        page.goto("https://pctsystem.cponline.cnipa.gov.cn")
        
        # Handle anti-detection measures

        page.wait_for_selector("#searchInput", timeout=10000)
        page.fill("#searchInput", application_number)
        page.click("#searchBtn")
        
        # Extract bibliographic data

        page.wait_for_load_state("networkidle")
        results["title"] = page.inner_text(".patent-title")
        results["applicant"] = page.inner_text(".applicant-name")
        results["abstract"] = page.inner_text(".abstract-content")
        
        browser.close()
    
    return results

def extract_pdf_from_result(page: Page, download_dir: str) -> str:
    with page.expect_download() as download_info:
        page.click("text=下载PDF")
    
    download = download_info.value
    file_path = f"{download_dir}/{download.suggested_filename}"
    download.save_as(file_path)
    
    return file_path

```

Playwright's synchronous API enables reliable PDF extraction and dynamic content scraping from JavaScript-heavy patent databases.

Source reference: [requirements.txt L9–L11](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L9)

### mammoth (≥1.6.0): HTML-to-Word Conversion

The `mammoth` package bridges markdown-based content with Word document generation. It converts HTML (rendered from markdown) into .docx format with preserved styling.

```python
import mammoth
from docx import Document
from io import BytesIO

def markdown_to_docx(html_content: str, style_map: str = None) -> Document:
    # Default style mapping for patent documents

    default_style_map = """
    p[style-name='Section Title'] => h1
    p[style-name='Subsection Title'] => h2
    p[style-name='Claim Text'] => p.claim
    """
    
    result = mammoth.convert_to_document(
        html_content,
        style_map=style_map or default_style_map
    )
    
    # Access warnings for debugging

    for warning in result.messages:
        print(f"Conversion warning: {warning}")
    
    return result.value  # Returns docx.Document object

def extract_raw_text_from_docx(file_path: str) -> str:
    with open(file_path, "rb") as f:
        result = mammoth.extract_raw_text(f)
    
    return result.value  # Plain text extraction

```

This appears in [`skills/patent-disclosure/tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/md_to_docx.py) for transforming structured markdown disclosures into submission-ready Word documents.

Source reference: [requirements.txt L11–L12](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L11)

### python-pptx (≥0.6.21): PowerPoint Presentation Generation

The `python-pptx` dependency enables automated creation of presentation materials from generated disclosure content, as implemented in [`skills/patent-disclosure/tools/pptx_to_md.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/pptx_to_md.py).

```python
from pptx import Presentation
from pptx.util import Inches, Pt
from pptx.enum.text import PP_ALIGN, MSO_ANCHOR
from pptx.dml.color import RGBColor

def create_disclosure_presentation(title: str, slides_data: List[Dict]) -> Presentation:
    prs = Presentation()
    prs.slide_width = Inches(13.333)
    prs.slide_height = Inches(7.5)
    
    # Title slide

    title_slide_layout = prs.slide_layouts[0]
    slide = prs.slides.add_slide(title_slide_layout)
    slide.shapes.title.text = title
    
    # Content slides

    for slide_data in slides_data:
        bullet_slide_layout = prs.slide_layouts[1]
        slide = prs.slides.add_slide(bullet_slide_layout)
        
        shapes = slide.shapes
        title_shape = shapes.title
        body_shape = shapes.placeholders[1]
        
        title_shape.text = slide_data["heading"]
        tf = body_shape.text_frame
        
        for item in slide_data["bullet_points"]:
            p = tf.add_paragraph()
            p.text = item
            p.level = 0
            p.font.size = Pt(18)
    
    return prs

def extract_pptx_to_markdown(pptx_path: str) -> str:
    """Reverse conversion: PowerPoint to markdown for archival."""
    prs = Presentation(pptx_path)
    md_lines = []
    
    for slide in prs.slides:
        if slide.shapes.title:
            md_lines.append(f"## {slide.shapes.title.text}\n")

        
        for shape in slide.shapes:
            if hasattr(shape, "text") and shape.text.strip():
                md_lines.append(f"- {shape.text.strip()}")
        
        md_lines.append("\n---\n")
    
    return "\n".join(md_lines)

```

Source reference: [requirements.txt L12–L13](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt#L12)

## Installation and Version Pinning

Install all core dependencies with standard pip commands:

```bash

# Install from requirements.txt

pip install -r requirements.txt

# Or install specific versions for reproducibility

pip install python-docx==1.1.0 latex2mathml==3.77.0 PyYAML==6.0 \
    playwright==1.40.0 mammoth==1.6.0 python-pptx==0.6.21

# Initialize Playwright browsers (one-time setup)

playwright install chromium

```

The version pinning in [`requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt) ensures compatibility across document generation workflows, particularly for `playwright` browser automation and `python-docx` formatting features.

## Optional Dependencies

Several sub-modules declare additional requirements for specialized functionality:

- **matplotlib** — Patent figure generation and claim diagram visualization
- **sqlite-vec** — Vector storage for semantic patent search
- **sentence-transformers** — Embedding models for claim similarity analysis
- **pymupdf** — Alternative PDF processing pipeline

These packages appear in sub-directory [`requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt) files and are not required for core patent-disclosure-skill operation as defined in the root [`requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt).

## Summary

- **python-docx (≥1.1.0)** generates and parses Word documents for patent disclosures and OA opinions
- **latex2mathml (≥3.77.0)** renders mathematical formulas in Word-compatible format
- **PyYAML (≥6.0)** loads configuration files and paradigm definitions
- **playwright (≥1.40.0)** automates CNIPA web searches and PDF extraction
- **mammoth (≥1.6.0)** converts HTML content to Word documents
- **python-pptx (≥0.6.21)** creates presentation slides from disclosure content

All dependencies are declared in [`requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/requirements.txt) with minimum version constraints, and their usage is demonstrated in [`skills/patent-disclosure/tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/md_to_docx.py), [`skills/patent-search/tools/browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/browser.py), and related utility modules.

## Frequently Asked Questions

### What is the minimum Python version for patent-disclosure-skill?

The repository targets Python 3.9 or higher based on type hint syntax and dependency requirements. Playwright 1.40.0+ requires Python 3.8+, while `python-docx` 1.1.0+ recommends Python 3.7+ with full 3.9+ feature support.

### Can I use patent-disclosure-skill without installing Playwright?

No, Playwright is a **core dependency** required for CNIPA patent search functionality. However, if you only need document generation features, you can modify imports to avoid the `skills/patent-search` module, though this requires source code changes.

### How does latex2mathml integrate with python-docx?

The `latex2mathml` package produces MathML strings that `python-docx` converts to OMML (Office Open XML Math) through internal XML namespace handling. In [`skills/patent-disclosure/tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-disclosure/tools/md_to_docx.py), this chain enables LaTeX formulas from technical disclosures to appear as native editable equations in generated Word documents.

### Are there security considerations for the Playwright dependency?

Yes. The [`skills/patent-search/tools/browser.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/browser.py) implementation launches headless Chromium with specific context settings. Production deployments should validate target URLs, implement request timeouts, and consider running browser automation in isolated containers due to the potential for untrusted web content execution.