How to Generate Patent Application Documents with the patent-disclosure-skill Toolkit

The patent-disclosure-skill toolkit converts Markdown-formatted patent sections—claims, description, abstract, and drawings—into submission-ready Microsoft Word documents using a specialized Python conversion engine that handles Chinese typography, mathematical formulas, and embedded images.

The handsomestWei/patent-disclosure-skill repository provides a dedicated patent-application skill that automates the final document preparation stage. This tool transforms structured Markdown files generated during the disclosure phase into properly formatted .docx files suitable for patent office submission, enforcing UTF-8 encoding and Chinese patent formatting standards throughout the pipeline.

Prerequisites and Required Files

Before running the generator, prepare a case directory containing the four standard patent document stems as Markdown files:

  • 权利要求书.md (Claims)
  • 说明书.md (Description)
  • 说明书摘要.md (Abstract)
  • 说明书附图.md (Drawings)

Place these files in a dedicated folder structure, for example outputs/patent-application/案例A/. The system processes any subset of these files, so you can generate documents incrementally if needed.

Step-by-Step Generation Process

Preparing the Source Directory

Create a case-specific folder and populate it with the Markdown files output from your disclosure workflow. The directory name becomes your case identifier. Ensure all Markdown files use UTF-8 encoding to prevent character corruption during conversion.

Running the Driver Script

Execute the conversion driver located at skills/patent-application/tools/emit_application_docx.py using the --dir flag to specify your case folder:

python tools/emit_application_docx.py --dir outputs/patent-application/案例A

The script performs three critical operations:

  1. UTF-8 enforcement – Calls stdio_utf8.ensure_utf8_stdio() to guarantee UTF-8 encoding for stdout, stderr, and child processes, preventing encoding errors when handling Chinese characters.

  2. File discovery – Locates each default stem (权利要求书, 说明书, 说明书摘要, 说明书附图) within the target directory by checking for corresponding .md files.

  3. Batch conversion – Invokes emit_one() for every existing Markdown file, which reads the source text and triggers the Word document generation pipeline.

Markdown to Word Conversion Engine

The core conversion logic resides in tools/md_to_docx.py within the convert_md_to_docx() function. This engine parses full Markdown syntax—including headings, lists, tables, images, and inline/block formulas—and constructs a python-docx Document object.

Image Handling: Embedded images are resized respecting maximum dimensions of approximately 5.5 inches by 8.2 inches to maintain proper page layout and patent office formatting requirements.

Mathematical Formulas: LaTeX expressions are first converted to OMML (Office Math Markup Language) for editable equations within Word. If OMML conversion fails, the system falls back to PNG image rendering or inserts the raw LaTeX text to ensure no data loss.

Typography: Styling normalizes to Chinese patent office defaults, specifically using 宋体 (SimSun) font at 10.5 pt size with standardized paragraph spacing.

The resulting Word document is saved adjacent to the source Markdown file with an identical filename stem but a .docx extension.

Result Reporting

Upon completion, the driver outputs machine-readable status lines for logging and automation:


DOCX: ok=1 path=outputs/patent-application/案例A/说明书.docx
APPLICATION_DOCX: ok=4 fail=0

This structured output enables CI/CD pipelines to verify successful document generation programmatically.

Programmatic API Usage

For integration into custom workflows, import the converter directly from your Python code:

from pathlib import Path
from tools.md_to_docx import convert_md_to_docx

md_path = Path("outputs/patent-application/案例A/说明书.md")
md_text = md_path.read_text(encoding="utf-8")
doc = convert_md_to_docx(md_text, base_dir=md_path.parent)
doc.save(md_path.with_suffix(".docx"))

This approach provides fine-grained control over the conversion process while maintaining the same formatting standards as the CLI tool.

Key Implementation Details

Isolated Architecture

The patent-application skill maintains a clean separation from the disclosure stage. As noted in the source comments of md_to_docx.py, the tool imports no external packages from the disclosure skill, ensuring the application generation phase remains self-contained and stable.

Formula Processing Pipeline

The conversion attempts LaTeX-to-OMML translation through the math_to_omml helper module. This preserves formula editability in Microsoft Word rather than rasterizing everything to static images, which is critical for patent examiners who may need to modify technical expressions during prosecution.

Encoding Guarantees

The stdio_utf8 utility ensures consistent UTF-8 handling across Windows and Unix-like systems, eliminating the "mojibake" character corruption commonly encountered when processing Chinese patent text on Western-localized operating systems.

Summary

  • Repository: handsomestWei/patent-disclosure-skill provides a dedicated patent-application skill for document finalization.
  • Input: Four Markdown files representing standard Chinese patent sections (claims, description, abstract, drawings).
  • Command: Run python tools/emit_application_docx.py --dir <case-folder> to generate .docx files.
  • Engine: convert_md_to_docx() in md_to_docx.py handles complex Markdown including tables, images, and LaTeX formulas.
  • Output: Word documents with Chinese patent formatting (SimSun 10.5pt) and machine-readable status reporting.
  • API: Direct Python import available for custom automation workflows.

Frequently Asked Questions

What file formats does the tool support for input and output?

The tool accepts Markdown (.md) files as input and generates Microsoft Word (.docx) documents as output. The conversion preserves complex formatting including headings, tables, mathematical formulas, and embedded images, producing files ready for direct submission to patent offices.

How does the tool handle mathematical formulas in patent claims?

The engine first attempts to convert LaTeX formulas to OMML (Office Math Markup Language), which creates editable equations in Word. If this conversion fails, it automatically falls back to PNG image rendering or raw LaTeX text insertion, ensuring technical accuracy while maximizing document usability for examiners.

Can I generate only specific sections of a patent application?

Yes. The driver script processes any subset of the four default stems (权利要求书, 说明书, 说明书摘要, 说明书附图). If only the claims and description files exist in your directory, the tool generates only those two Word documents and reports the successful conversions in its status output.

What dependencies are required to run the conversion?

The tool requires python-docx for Word document generation and optionally matplotlib for PNG fallback rendering of formulas. Install dependencies via skills/patent-application/tools/requirements.txt. The code enforces UTF-8 I/O automatically through the included stdio_utf8 utility, requiring no additional system configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →