# How to Generate Patent Application Documents with the patent-disclosure-skill Toolkit

> Generate patent application documents from Markdown. The patent-disclosure-skill toolkit converts text, formulas, and images into Word docs with a Python engine.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-08

---

**The patent-disclosure-skill toolkit converts Markdown-formatted patent sections—claims, description, abstract, and drawings—into submission-ready Microsoft Word documents using a specialized Python conversion engine that handles Chinese typography, mathematical formulas, and embedded images.**

The `handsomestWei/patent-disclosure-skill` repository provides a dedicated *patent-application* skill that automates the final document preparation stage. This tool transforms structured Markdown files generated during the disclosure phase into properly formatted `.docx` files suitable for patent office submission, enforcing UTF-8 encoding and Chinese patent formatting standards throughout the pipeline.

## Prerequisites and Required Files

Before running the generator, prepare a case directory containing the four standard patent document stems as Markdown files:

- `权利要求书.md` (Claims)
- `说明书.md` (Description)
- `说明书摘要.md` (Abstract)
- `说明书附图.md` (Drawings)

Place these files in a dedicated folder structure, for example `outputs/patent-application/案例A/`. The system processes any subset of these files, so you can generate documents incrementally if needed.

## Step-by-Step Generation Process

### Preparing the Source Directory

Create a case-specific folder and populate it with the Markdown files output from your disclosure workflow. The directory name becomes your case identifier. Ensure all Markdown files use UTF-8 encoding to prevent character corruption during conversion.

### Running the Driver Script

Execute the conversion driver located at [`skills/patent-application/tools/emit_application_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-application/tools/emit_application_docx.py) using the `--dir` flag to specify your case folder:

```bash
python tools/emit_application_docx.py --dir outputs/patent-application/案例A

```

The script performs three critical operations:

1. **UTF-8 enforcement** – Calls `stdio_utf8.ensure_utf8_stdio()` to guarantee UTF-8 encoding for stdout, stderr, and child processes, preventing encoding errors when handling Chinese characters.

2. **File discovery** – Locates each default stem (`权利要求书`, `说明书`, `说明书摘要`, `说明书附图`) within the target directory by checking for corresponding `.md` files.

3. **Batch conversion** – Invokes `emit_one()` for every existing Markdown file, which reads the source text and triggers the Word document generation pipeline.

### Markdown to Word Conversion Engine

The core conversion logic resides in [`tools/md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/md_to_docx.py) within the `convert_md_to_docx()` function. This engine parses full Markdown syntax—including headings, lists, tables, images, and inline/block formulas—and constructs a `python-docx` Document object.

**Image Handling**: Embedded images are resized respecting maximum dimensions of approximately 5.5 inches by 8.2 inches to maintain proper page layout and patent office formatting requirements.

**Mathematical Formulas**: LaTeX expressions are first converted to **OMML** (Office Math Markup Language) for editable equations within Word. If OMML conversion fails, the system falls back to PNG image rendering or inserts the raw LaTeX text to ensure no data loss.

**Typography**: Styling normalizes to Chinese patent office defaults, specifically using **宋体** (SimSun) font at **10.5 pt** size with standardized paragraph spacing.

The resulting Word document is saved adjacent to the source Markdown file with an identical filename stem but a `.docx` extension.

### Result Reporting

Upon completion, the driver outputs machine-readable status lines for logging and automation:

```

DOCX: ok=1 path=outputs/patent-application/案例A/说明书.docx
APPLICATION_DOCX: ok=4 fail=0

```

This structured output enables CI/CD pipelines to verify successful document generation programmatically.

## Programmatic API Usage

For integration into custom workflows, import the converter directly from your Python code:

```python
from pathlib import Path
from tools.md_to_docx import convert_md_to_docx

md_path = Path("outputs/patent-application/案例A/说明书.md")
md_text = md_path.read_text(encoding="utf-8")
doc = convert_md_to_docx(md_text, base_dir=md_path.parent)
doc.save(md_path.with_suffix(".docx"))

```

This approach provides fine-grained control over the conversion process while maintaining the same formatting standards as the CLI tool.

## Key Implementation Details

### Isolated Architecture

The patent-application skill maintains a clean separation from the disclosure stage. As noted in the source comments of [`md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/md_to_docx.py), the tool imports no external packages from the disclosure skill, ensuring the application generation phase remains self-contained and stable.

### Formula Processing Pipeline

The conversion attempts LaTeX-to-OMML translation through the `math_to_omml` helper module. This preserves formula editability in Microsoft Word rather than rasterizing everything to static images, which is critical for patent examiners who may need to modify technical expressions during prosecution.

### Encoding Guarantees

The `stdio_utf8` utility ensures consistent UTF-8 handling across Windows and Unix-like systems, eliminating the "mojibake" character corruption commonly encountered when processing Chinese patent text on Western-localized operating systems.

## Summary

- **Repository**: `handsomestWei/patent-disclosure-skill` provides a dedicated patent-application skill for document finalization.
- **Input**: Four Markdown files representing standard Chinese patent sections (claims, description, abstract, drawings).
- **Command**: Run `python tools/emit_application_docx.py --dir <case-folder>` to generate `.docx` files.
- **Engine**: `convert_md_to_docx()` in [`md_to_docx.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/md_to_docx.py) handles complex Markdown including tables, images, and LaTeX formulas.
- **Output**: Word documents with Chinese patent formatting (SimSun 10.5pt) and machine-readable status reporting.
- **API**: Direct Python import available for custom automation workflows.

## Frequently Asked Questions

### What file formats does the tool support for input and output?

The tool accepts **Markdown (.md)** files as input and generates **Microsoft Word (.docx)** documents as output. The conversion preserves complex formatting including headings, tables, mathematical formulas, and embedded images, producing files ready for direct submission to patent offices.

### How does the tool handle mathematical formulas in patent claims?

The engine first attempts to convert LaTeX formulas to **OMML** (Office Math Markup Language), which creates editable equations in Word. If this conversion fails, it automatically falls back to PNG image rendering or raw LaTeX text insertion, ensuring technical accuracy while maximizing document usability for examiners.

### Can I generate only specific sections of a patent application?

Yes. The driver script processes any subset of the four default stems (`权利要求书`, `说明书`, `说明书摘要`, `说明书附图`). If only the claims and description files exist in your directory, the tool generates only those two Word documents and reports the successful conversions in its status output.

### What dependencies are required to run the conversion?

The tool requires `python-docx` for Word document generation and optionally `matplotlib` for PNG fallback rendering of formulas. Install dependencies via [`skills/patent-application/tools/requirements.txt`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-application/tools/requirements.txt). The code enforces UTF-8 I/O automatically through the included `stdio_utf8` utility, requiring no additional system configuration.