Difference Between Mermaid, Graphviz, and PlantUML Text Representations in TextFlow

TextFlow treats Mermaid, Graphviz, and PlantUML as alternative text-based diagram formats, but only Mermaid offers native serialization through Flowchart.to_mermaid(), while Graphviz and PlantUML rely entirely on LLM generation and fence-based extraction utilities.

The junyiye/textflow repository provides a flexible framework for converting visual flowcharts into machine-readable text representations. Understanding the difference between Mermaid, Graphviz, and PlantUML text representations in TextFlow is essential for selecting the right output format for your documentation pipeline, as each option varies significantly in native support, extraction methodology, and downstream compatibility.

Native Serialization vs. LLM-Generated Output

Mermaid: Built-in Object Serialization

Mermaid is the only format in TextFlow that provides native serialization from internal data structures. The Flowchart class defined in src/flowchart.py implements a to_mermaid() method that converts the internal node and edge representation directly into valid Mermaid syntax. This ensures syntactically correct output without requiring LLM intervention for structural conversion.

Graphviz and PlantUML: LLM-Dependent Generation

Unlike Mermaid, Graphviz (DOT) and PlantUML lack native serializers in the TextFlow codebase. When you specify --output_type graphviz or --output_type plantuml, the system relies entirely on the Vision Language Model (VLM) to generate syntactically correct diagram code. The LLM must produce properly fenced code blocks (dot or plantuml), which TextFlow then extracts using specialized utility functions.

Extraction and Processing Pipeline

TextFlow handles all three formats through a unified extraction pipeline defined in src/utils.py. The system identifies the diagram type by detecting specific fence tags within the LLM response:

  • extract_mermaid_code (lines 47-55): Extracts content from mermaid fences
  • extract_graphviz_code (lines 60-68): Extracts content from dot fences
  • extract_plantuml_code (lines 73-81): Extracts content from plantuml fences

The generic extract_representation() function acts as a dispatcher, automatically detecting which fence tag is present and invoking the appropriate specialized extractor. This design allows the downstream reasoner.py module (lines 38-39) to accept any of the three input formats interchangeably for further processing.

CLI Configuration and Usage

The src/textualizer.py module provides command-line interface controls for selecting your desired output format. By default, TextFlow uses Mermaid representation:


# Default Mermaid output (lines 33-37 in textualizer.py)

python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type mermaid

To generate alternative formats, explicitly specify the output type:


# Graphviz DOT format

python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type graphviz

# PlantUML format  

python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type plantuml

All three commands store the extracted diagram string under output/flowvqa/<type>/<textualizer>.json, maintaining consistent file organization regardless of the representation format.

Programmatic Conversion Examples

When working with internal Flowchart objects, only Mermaid provides direct serialization:

from src.flowchart import Flowchart

# Create and populate flowchart

fc = Flowchart()

# ... add nodes and edges ...

# Native Mermaid conversion

mermaid_src = fc.to_mermaid()
print(mermaid_src)

For Graphviz or PlantUML, you must prompt an LLM to generate the representation and then extract it using the utility functions:

from src.utils import extract_representation

# Simulated LLM response containing PlantUML

response = """
Here is the diagram:

```plantuml
@startuml
A -> B
B -> C
@enduml

"""

diagram = extract_representation(response)


## Summary

- **Mermaid** is the default and only format with native serialization support via `Flowchart.to_mermaid()` in [`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py), making it the most reliable choice for programmatic generation.

- **Graphviz and PlantUML** lack internal serializers; TextFlow relies on LLM generation and extracts the resulting code from fenced blocks using `extract_graphviz_code` and `extract_plantuml_code` in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py).

- All three formats are supported by the CLI `--output_type` argument in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) and can be processed downstream by [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), but only Mermaid guarantees syntactic correctness without LLM intervention.

## Frequently Asked Questions

### Can I convert a TextFlow Flowchart object directly to Graphviz DOT format?

No. Unlike Mermaid, TextFlow does not implement a `to_graphviz()` method in the `Flowchart` class. To obtain Graphviz output, you must use the `--output_type graphviz` CLI flag or prompt an LLM to generate DOT syntax, then extract it using `utils.extract_graphviz_code()`.

### Why does TextFlow default to Mermaid instead of Graphviz or PlantUML?

Mermaid is the default because it is the only format that provides **native serialization** through `Flowchart.to_mermaid()`. This ensures that the internal flowchart structure can be converted to valid text representation without relying on LLM generation, reducing the risk of syntax errors and improving reproducibility.

### How does TextFlow validate the syntax of generated PlantUML or Graphviz code?

TextFlow does not perform native syntax validation for PlantUML or Graphviz code. The system relies on the LLM to produce syntactically correct output within the appropriate fenced code blocks (````dot```` or ````plantuml````). The [`utils.py`](https://github.com/junyiye/textflow/blob/main/utils.py) module only handles extraction via `extract_graphviz_code` and `extract_plantuml_code`, passing the raw string to downstream components without parsing validation.

### Which text representation should I use for GitHub documentation?

Use **Mermaid** for GitHub documentation. GitHub's Markdown renderer natively supports Mermaid diagrams in fenced code blocks, allowing immediate visualization without external tooling. While Graphviz and PlantUML require additional rendering pipelines or extensions, Mermaid diagrams generated by TextFlow can be viewed directly in pull requests, issues, and README files on GitHub.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →