How the Invention Disclosure Builder Incorporates Mermaid Flowcharts into Word Documents
The invention disclosure builder converts Mermaid flowcharts in Markdown to embedded PNG images in Word documents through a two-step pipeline: first rendering diagrams to PNG via Playwright with hidden HTML comments as markers, then parsing those comments during Markdown-to-Word conversion to embed the images.
The handsomestWei/patent-disclosure-skill repository provides a specialized toolchain for generating patent invention disclosures from Markdown source files. A key capability is the seamless integration of Mermaid flowcharts—textual diagram definitions—into final Microsoft Word (.docx) outputs. This article explains the exact mechanism, tracing the code path from mermaid_render.py through md_to_docx.py.
The Two-Step Rendering Architecture
The builder separates concerns into distinct phases: diagram generation and document assembly. This design keeps the Markdown source human-readable while ensuring polished visual output in Word.
Step 1: Render Mermaid Blocks to PNG Images
The tools/shared/mermaid_render.py module handles the heavy lifting of converting Mermaid syntax into raster images. It operates on a draft Markdown file and produces an enriched version ready for Word conversion.
Detecting Mermaid Fences
The script scans input Markdown for fenced code blocks beginning with ```mermaid. The rendering logic lives in the render_one function (approximately lines 32–48), which orchestrates the browser-based conversion.
Headless Browser Rendering with Playwright
Mermaid diagrams require a JavaScript execution environment. The _PlaywrightMermaid context manager (lines 82–110) encapsulates this:
from tools.shared.mermaid_render import render_markdown_mermaid
# Process draft.md, writing PNGs and enriched Markdown
render_markdown_mermaid(
input_path="draft.md",
output_path="disclosure.md",
image_dir="./images"
)
Inside the browser, the script:
- Loads a staging HTML template (
_STAGE_HTML) with an empty#stageelement - Injects the vendored
mermaid.min.jslibrary - Executes
_RENDER_JS(lines 60–74), which callsmermaid.render()to generate SVG - Captures a screenshot via
page.locator("#stage").screenshot()(line 47)
The PNG output is written to a configurable directory, typically ./images/mermaid-{n}.png.
Inserting Hidden HTML Comments
Crucially, the original Mermaid fence remains intact in the Markdown. After each fenced block, the renderer appends a hidden HTML comment that serves as a metadata marker:
```mermaid
graph TB
A[Invention Concept] --> B{Novel?}
B -->|Yes| C[Proceed to Claims]
B -->|No| D[Document Prior Art]
The comment format is governed by `_MERMAID_HIDDEN_COMMENT_RE` (lines 52–54), which matches the pattern `<!--  -->`. This convention enables lossless round-tripping: the Markdown stays readable, while downstream processors can locate and embed the corresponding images.
### Step 2: Convert Enriched Markdown to Word
The `tools/shared/md_to_docx.py` module completes the pipeline by transforming the enriched Markdown into a `.docx` file, handling the hidden comment markers to embed PNGs.
#### Parsing Hidden Image Comments
The converter uses `_HIDDEN_MD_IMAGE_COMMENT_RE` (lines 57–58) to detect the special HTML comments inserted in step 1. The helper function `_parse_hidden_image_comment()` (lines 93–100) extracts:
- **Alt text**: The description inside `![...]` (e.g., "图示" meaning "illustration")
- **Image path**: The relative path inside `(...)`
#### Embedding Images with python-docx
Once extracted, the PNG is embedded into the Word document using `python-docx` primitives. The `_embed_from_image_ref` helper and related functions manage:
- Loading the image file from the extracted path
- Sizing and positioning within the document flow
- Maintaining the association between the Mermaid code block (preserved as literal text) and its visual representation
### Complete Pipeline Execution
The full workflow from draft to final document:
```bash
# Step 1: Generate PNGs and enriched Markdown
python tools/shared/mermaid_render.py \
--input draft.md \
--output disclosure.md \
--image-dir ./images
# Step 2: Produce Word document with embedded diagrams
python tools/shared/md_to_docx.py \
--input disclosure.md \
--output disclosure.docx
After execution, disclosure.docx contains both the original Mermaid source (for archival/review purposes) and the rendered flowchart image positioned at the corresponding location.
Shared Infrastructure: Browser Management
Both steps rely on tools/shared/browser.py, which provides vendor-agnostic Chromium launching via Playwright. Key utilities include:
launch_chromium(): Configures headless browser instanceplaywright_installed(): Verifies environment readiness
This centralization ensures consistent browser behavior across rendering and testing workflows.
Testing the Integration
The repository includes tests/shared/test_mermaid_browser.py, which validates end-to-end Mermaid rendering. This test confirms that:
- Mermaid fences are correctly identified
- PNG files are generated with valid image data
- Hidden comments follow the expected format
Running these tests ensures the two-step pipeline remains functional across environment changes.
Configuration and Customization
Several parameters control the rendering behavior:
| Parameter | Location | Effect |
|---|---|---|
--image-dir / image_dir |
mermaid_render.py CLI |
Output directory for PNG files |
--mermaid-theme |
Mermaid config | Visual styling of diagrams (default, dark, forest, neutral) |
| Browser timeout | _PlaywrightMermaid |
Maximum wait for diagram rendering |
Advanced users can modify _STAGE_HTML or _RENDER_JS in mermaid_render.py to inject custom CSS or Mermaid configuration options.
Summary
- Separation of concerns: Diagram rendering (
mermaid_render.py) and document assembly (md_to_docx.py) are distinct, composable steps - Hidden HTML comments: Serve as lossless metadata markers linking Mermaid source blocks to their PNG renderings
- Playwright-based rendering: Provides accurate, browser-consistent diagram generation without external dependencies like Node.js CLI tools
- Preserved source: Original Mermaid fences remain in the Word document, enabling future editing and audit trails
Frequently Asked Questions
How does the builder handle Mermaid syntax errors?
mermaid_render.py delegates error handling to the Mermaid.js library running in Playwright. If mermaid.render() fails, the JavaScript exception propagates through the browser context, causing render_one to raise a runtime error. The original Markdown fence is not modified, so the user can correct the syntax and re-run.
Can I customize the image format or resolution?
Currently, browser.py and mermaid_render.py generate PNG via Playwright's screenshot API. The resolution is determined by the browser viewport size configured in _STAGE_HTML. To change formats (e.g., SVG embedding), you would need to modify render_one to capture SVG source directly and adjust md_to_docx.py to handle vector image embedding.
Why use hidden HTML comments instead of standard Markdown image syntax?
Standard Markdown images () would render visibly in HTML previews and confuse the document structure. Hidden HTML comments are invisible in rendered output but parseable by md_to_docx.py. This preserves a clean reading experience while enabling the Word converter to locate and substitute the correct visual assets.
Is Playwright required for all users of the disclosure builder?
Yes, Playwright and its Chromium browser are required dependencies for mermaid_render.py. The browser.py module provides playwright_installed() to verify the environment. Users without Playwright can still run md_to_docx.py on pre-rendered Markdown files containing the hidden comment markers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →