How Invention Formulas Are Processed and Rendered in Word: A Technical Deep Dive

The repository converts LaTeX-style invention formulas embedded in Markdown into properly sized PNG images via matplotlib-mathtext, then embeds them into Word documents as either inline or centered block elements while preserving mathematical fidelity.

The handsomestWei/patent-disclosure-skill repository implements a specialized pipeline for converting patent disclosure documents containing mathematical formulas into Microsoft Word format. This system bridges the gap between Markdown-based drafting and professional Word document requirements by handling LaTeX syntax detection, image rendering, and precise document layout automatically.

The Three-Stage Processing Pipeline

The invention formula workflow operates through distinct detection, rendering, and embedding phases coordinated across two primary modules.

Stage 1: LaTeX Detection and Extraction

The process begins in tools/shared/math_render.py, where the parser identifies mathematical expressions using regex patterns. The system recognizes inline formulas delimited by $...$ or \(...\) syntax, and block formulas marked by $$...$$ or \[ ... \] delimiters.

The _INLINE_RE pattern and iter_inline_paren_spans() function scan the Markdown source to locate formula boundaries. Upon detection, each formula is normalized—converting commands like \geqslant to \geq—to ensure compatibility with the rendering engine. Failed detections are logged silently to prevent document corruption.

Stage 2: PNG Rendering via matplotlib-mathtext

Once extracted, formulas undergo rendering through the render_latex_to_png() function. This implementation leverages mathtext.math_to_image() (lines 31–42) from the matplotlib library to convert LaTeX strings into high-resolution PNG assets.

The _try_render() helper manages success and failure states, generating sequential filenames such as inline_001.png or eq_001.png within a configurable math_figures directory. This stage produces mathematically accurate raster images with transparent backgrounds suitable for Word integration.

Stage 3: Word Document Embedding

The final transformation occurs in tools/shared/md_to_docx.py, which processes the augmented Markdown containing hidden HTML comments. These comments—formatted as <!-- ![公式·行内](math_figures/inline_001.png) -->—serve as embedded references for the Word converter.

The _formula_image_kind() function analyzes the alt text (e.g., "公式·行内" for inline, "公式·块" for block) and file paths to determine insertion semantics. Inline formulas are embedded using _embed_picture_inline() with constrained heights (_FORMULA_INLINE_MAX_H_IN), while block formulas generate new paragraphs aligned with WD_ALIGN_PARAGRAPH.CENTER including proper spacing before and after.

Command-Line Workflow Implementation

Executing the complete pipeline requires sequential tool invocations that transform raw Markdown into publication-ready Word documents.

First, render all LaTeX formulas to PNG assets and generate the intermediate Markdown with hidden image references:

python tools/shared/math_render.py -i patent.md -o patent_with_imgs.md --assets-dir math_figures

Then convert the augmented Markdown into a Word document with embedded mathematics:

python tools/shared/md_to_docx.py -i patent_with_imgs.md -o patent.docx --math-render

Input files containing standard LaTeX syntax process automatically:

The energy-mass relation is $E = mc^2$.

$$
\int_{0}^{\infty} e^{-x^2}\,dx = \frac{\sqrt{\pi}}{2}
$$

After processing, inline formulas appear seamlessly within text lines, while block formulas render as centered, standalone equations with professional typography spacing.

Inline vs. Block Formula Handling

The system distinguishes between formula presentation modes through metadata encoded in hidden comments and path analysis.

Inline formulas maintain typographical integration with surrounding text. The _embed_from_image_ref() function loads PNG assets, calculates pixel dimensions via _image_pixel_size(), and computes display measurements using _fit_image_display_inches(). These formulas insert into the current paragraph run without interrupting text flow.

Block formulas trigger paragraph-level insertion. When _formula_image_kind() returns "block", the converter creates a new paragraph with center alignment, establishing visual separation appropriate for displayed equations in formal patent documents.

Fallback Strategies: OMML and Raw Text Preservation

When hidden image comments are absent from the Markdown source—such as when the math renderer was not run—the system implements hierarchical fallback mechanisms.

The _try_append_omml() function attempts conversion to Office Math Markup Language (OMML) using utilities from tools/shared/math_to_omml.py. This produces native Word equation objects that remain editable.

If OMML generation fails, the pipeline falls back to PNG generation on-the-fly. Should image embedding also fail, the system preserves the original LaTeX source text via _add_math_fallback_block(), ensuring no mathematical content is lost during conversion, albeit without rendering.

Summary

  • Detection occurs in tools/shared/math_render.py using regex patterns _INLINE_RE to identify $...$ and $$...$$ syntax.
  • Rendering leverages matplotlib.mathtext.math_to_image() via render_latex_to_png() to create PNG assets in configurable directories.
  • Embedding logic in tools/shared/md_to_docx.py uses _formula_image_kind() to distinguish inline insertion from centered block paragraphs.
  • Fallback chains prioritize OMML native equations, then PNG images, then raw LaTeX text to prevent data loss.
  • Style metadata stored in tools/shared/formula_paradigms.py informs rendering parameters and document-specific formatting requirements.

Frequently Asked Questions

What LaTeX delimiters are supported for invention formulas?

The system recognizes four delimiter patterns: single dollar signs $...$ and backslash-parentheses \(...\) for inline formulas, and double dollar signs $$...$$ or backslash-brackets \[ ... \] for block-level equations. These patterns are detected by the iter_inline_paren_spans() generator and the _INLINE_RE regular expression defined in tools/shared/math_render.py.

How does the converter decide between inline and block rendering?

The _formula_image_kind() function in md_to_docx.py examines the hidden HTML comment's alt attribute—specifically looking for "公式·行内" (inline) or "公式·块" (block) markers—and analyzes the image file path. Inline formulas embed directly into the active paragraph, while block formulas generate new centered paragraphs using WD_ALIGN_PARAGRAPH.CENTER alignment constants.

What occurs when matplotlib fails to render a complex formula?

The _try_render() wrapper catches rendering exceptions and records failure states. If PNG generation fails during the initial math_render.py phase, that specific formula remains unconverted in the Markdown. During the Word conversion phase, if OMML conversion via _try_append_omml() also fails, the system invokes _add_math_fallback_block() to insert the raw LaTeX source text, ensuring mathematical content remains accessible even without visual rendering.

Can the output directory for formula images be customized?

Yes. Both tools accept an --assets-dir parameter to specify the destination folder for PNG files. The math_render.py script writes rendered images to this directory and generates corresponding hidden comments pointing to these paths. The md_to_docx.py converter then resolves these relative paths when embedding images into the final Word document.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →