How to Convert Markdown to DOCX for Patent Documents Using Python
The patent-disclosure-skill employs a deterministic line-by-line parser in md_to_docx.py to map Markdown syntax to python-docx constructs, with specialized handlers for patent-specific elements like mathematical formulas, embedded images, and Chinese typography standards.
Converting technical patent drafts from Markdown to Microsoft Word requires precise control over document structure, mathematical notation, and localized formatting. The handsomestWei/patent-disclosure-skill repository provides a specialized conversion tool that bridges this gap through a deterministic parsing pipeline. This implementation processes patent documents line-by-line to generate .docx files that maintain visual fidelity while supporting editable equations and standardized Chinese patent office formatting.
The Core Conversion Pipeline
Entry Point and CLI Configuration
The conversion begins in main() within skills/patent-oa/tools/md_to_docx.py (mirrored in skills/patent-disclosure/tools/md_to_docx.py), which constructs an argparse.ArgumentParser to handle input sources and rendering options. The parser accepts flags such as --no-omml to disable Office Math XML conversion and --math-render to generate PNG assets for LaTeX formulas. The function resolves the base directory for relative image paths and reads the Markdown source from either a specified file or STDIN.
Document Initialization
Before parsing begins, the converter creates a new python-docx Document instance and configures the default "Normal" paragraph style. The style applies the Chinese font 宋体 (SimSun) at 10.5 pt to ensure compliance with standard patent document formatting requirements. This initialization establishes the typographic foundation that persists throughout the generated document.
Line-by-Line Parsing Strategy
The core algorithm splits the Markdown source into individual lines and processes them within a state-managed loop. Blank lines trigger the flushing of the current paragraph buffer, ensuring proper spacing between document elements. Special block markers—including fenced code delimiters (```), mathematical boundaries (\[/\] or $$), ATX headings, list indicators, table dividers, and blockquotes—activate dedicated handler functions that construct the corresponding Word document elements.
Block-Level Element Handling
Code Fences and Horizontal Rules
Fenced code blocks are rendered as monospaced paragraphs using the Consolas font via the _add_code_block() function. Horizontal rules (thematic breaks) are converted into decorative lines composed of Unicode box-drawing characters ("─") rendered in light gray, providing visual separation without arbitrary shape insertion.
Tables and List Management
Simple GFM-style tables are parsed by _parse_table_row() and constructed through _add_table(), which applies the "Table Grid" style and recursively injects cell content. For lists, the _add_list_item() function manages both unordered (-, *, +) and ordered (1.) markers. Each ordered list receives a fresh numbering instance to ensure enumeration restarts correctly after interrupting elements like headings or tables.
Rich Content Processing
Image Embedding and Sizing Constraints
The _embed_from_image_ref() function resolves image paths relative to the specified base directory and categorizes assets as formulas, diagrams, or generic images. The converter reads pixel dimensions directly from PNG, GIF, and JPEG files using _image_pixel_size() without external dependencies like Pillow, converting measurements using a fixed 220 DPI. Sizing constraints vary by type: inline formula images are limited to 0.22 inches height, block formulas to 0.36 inches height and 5.5 inches width, while generic images max out at 5.5 inches width and 8.2 inches height.
Mathematical Formula Conversion
LaTeX fragments undergo a prioritized conversion pipeline. First, math_to_omml.try_latex_to_omml() attempts to generate editable Office Math XML (OMML) via the _try_append_omml() method. When OMML conversion fails, the system falls back to PNG rendering through math_render.render_markdown_math() if the --math-render flag is active. As a final resort, the raw LaTeX source is preserved in a monospaced code block via _add_math_fallback_block(). Conversion statistics are tracked in MathOutcomeStats and reported to STDERR upon completion.
Command-Line Usage Examples
Basic conversion with image path resolution:
python skills/patent-oa/tools/md_to_docx.py \
-i patent_specification.md \
-o patent_specification.docx \
--base-dir .
Programmatic integration within Python applications:
from pathlib import Path
from skills.patent_oa.tools.md_to_docx import convert_md_to_docx
md_path = Path("patent_specification.md")
doc = convert_md_to_docx(
md_path.read_text(encoding="utf-8"),
base_dir=md_path.parent,
image_max_w_in=5.5,
image_max_h_in=8.2,
prefer_omml=True,
)
doc.save("patent_specification.docx")
Rendering LaTeX to PNG for complex formulas:
python skills/patent-oa/tools/md_to_docx.py \
-i patent_specification.md \
-o patent_specification.docx \
--math-render
Summary
- The conversion pipeline resides in
skills/patent-oa/tools/md_to_docx.py(with an identical implementation inskills/patent-disclosure/tools/md_to_docx.py), implementing a line-by-line parser that maps Markdown constructs topython-docxobjects. - Document initialization sets 宋体 (SimSun) 10.5pt as the default font to meet Chinese patent formatting standards, while code blocks use Consolas.
- Images are processed with DPI-aware sizing constraints, distinguishing between inline formulas (0.22in height), block formulas (0.36in height, 5.5in width), and general diagrams (5.5in width, 8.2in height).
- Mathematical content attempts OMML conversion first, falling back to PNG rendering via
math_render.pyor plain text based on configuration flags. - The tool supports both CLI usage with
argparseand programmatic integration via theconvert_md_to_docx()function.
Frequently Asked Questions
How does the skill handle mathematical formulas in Markdown?
The converter prioritizes editable Office Math XML (OMML) using math_to_omml.try_latex_to_omml(). If LaTeX parsing fails, it checks for the --math-render flag to generate PNG images via math_render.render_markdown_math(), or defaults to displaying the raw LaTeX in a monospaced code block as a last resort.
What image formats are supported for embedding?
The _image_pixel_size() function natively supports PNG, GIF, and JPEG formats by reading file headers directly without Pillow dependency. Embedded images respect patent-specific dimension constraints based on classification as formulas or diagrams.
Can I disable OMML conversion and use PNG rendering instead?
Yes. Passing the --no-omml flag disables OMML generation, while --math-render enables PNG rendering for all mathematical expressions. If both options are omitted and OMML fails, the system falls back to plain text representation.
How are patent document fonts configured?
The converter explicitly sets the document's default "Normal" style to use 宋体 (SimSun) at 10.5 points during initialization in skills/patent-oa/tools/md_to_docx.py, while code blocks use Consolas. These defaults align with standard Chinese patent office submission requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →