Difference Between Mermaid, Graphviz, and PlantUML Text Representations in TextFlow
TextFlow treats Mermaid, Graphviz, and PlantUML as alternative text-based diagram formats, but only Mermaid offers native serialization through Flowchart.to_mermaid(), while Graphviz and PlantUML rely entirely on LLM generation and fence-based extraction utilities.
The junyiye/textflow repository provides a flexible framework for converting visual flowcharts into machine-readable text representations. Understanding the difference between Mermaid, Graphviz, and PlantUML text representations in TextFlow is essential for selecting the right output format for your documentation pipeline, as each option varies significantly in native support, extraction methodology, and downstream compatibility.
Native Serialization vs. LLM-Generated Output
Mermaid: Built-in Object Serialization
Mermaid is the only format in TextFlow that provides native serialization from internal data structures. The Flowchart class defined in src/flowchart.py implements a to_mermaid() method that converts the internal node and edge representation directly into valid Mermaid syntax. This ensures syntactically correct output without requiring LLM intervention for structural conversion.
Graphviz and PlantUML: LLM-Dependent Generation
Unlike Mermaid, Graphviz (DOT) and PlantUML lack native serializers in the TextFlow codebase. When you specify --output_type graphviz or --output_type plantuml, the system relies entirely on the Vision Language Model (VLM) to generate syntactically correct diagram code. The LLM must produce properly fenced code blocks (dot or plantuml), which TextFlow then extracts using specialized utility functions.
Extraction and Processing Pipeline
TextFlow handles all three formats through a unified extraction pipeline defined in src/utils.py. The system identifies the diagram type by detecting specific fence tags within the LLM response:
extract_mermaid_code(lines 47-55): Extracts content frommermaidfencesextract_graphviz_code(lines 60-68): Extracts content fromdotfencesextract_plantuml_code(lines 73-81): Extracts content fromplantumlfences
The generic extract_representation() function acts as a dispatcher, automatically detecting which fence tag is present and invoking the appropriate specialized extractor. This design allows the downstream reasoner.py module (lines 38-39) to accept any of the three input formats interchangeably for further processing.
CLI Configuration and Usage
The src/textualizer.py module provides command-line interface controls for selecting your desired output format. By default, TextFlow uses Mermaid representation:
# Default Mermaid output (lines 33-37 in textualizer.py)
python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type mermaid
To generate alternative formats, explicitly specify the output type:
# Graphviz DOT format
python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type graphviz
# PlantUML format
python -m src.textualizer --dataset flowvqa --textualizer Qwen2-VL-7B --output_type plantuml
All three commands store the extracted diagram string under output/flowvqa/<type>/<textualizer>.json, maintaining consistent file organization regardless of the representation format.
Programmatic Conversion Examples
When working with internal Flowchart objects, only Mermaid provides direct serialization:
from src.flowchart import Flowchart
# Create and populate flowchart
fc = Flowchart()
# ... add nodes and edges ...
# Native Mermaid conversion
mermaid_src = fc.to_mermaid()
print(mermaid_src)
For Graphviz or PlantUML, you must prompt an LLM to generate the representation and then extract it using the utility functions:
from src.utils import extract_representation
# Simulated LLM response containing PlantUML
response = """
Here is the diagram:
```plantuml
@startuml
A -> B
B -> C
@enduml
"""
diagram = extract_representation(response)
## Summary
- **Mermaid** is the default and only format with native serialization support via `Flowchart.to_mermaid()` in [`src/flowchart.py`](https://github.com/junyiye/textflow/blob/main/src/flowchart.py), making it the most reliable choice for programmatic generation.
- **Graphviz and PlantUML** lack internal serializers; TextFlow relies on LLM generation and extracts the resulting code from fenced blocks using `extract_graphviz_code` and `extract_plantuml_code` in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py).
- All three formats are supported by the CLI `--output_type` argument in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) and can be processed downstream by [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), but only Mermaid guarantees syntactic correctness without LLM intervention.
## Frequently Asked Questions
### Can I convert a TextFlow Flowchart object directly to Graphviz DOT format?
No. Unlike Mermaid, TextFlow does not implement a `to_graphviz()` method in the `Flowchart` class. To obtain Graphviz output, you must use the `--output_type graphviz` CLI flag or prompt an LLM to generate DOT syntax, then extract it using `utils.extract_graphviz_code()`.
### Why does TextFlow default to Mermaid instead of Graphviz or PlantUML?
Mermaid is the default because it is the only format that provides **native serialization** through `Flowchart.to_mermaid()`. This ensures that the internal flowchart structure can be converted to valid text representation without relying on LLM generation, reducing the risk of syntax errors and improving reproducibility.
### How does TextFlow validate the syntax of generated PlantUML or Graphviz code?
TextFlow does not perform native syntax validation for PlantUML or Graphviz code. The system relies on the LLM to produce syntactically correct output within the appropriate fenced code blocks (````dot```` or ````plantuml````). The [`utils.py`](https://github.com/junyiye/textflow/blob/main/utils.py) module only handles extraction via `extract_graphviz_code` and `extract_plantuml_code`, passing the raw string to downstream components without parsing validation.
### Which text representation should I use for GitHub documentation?
Use **Mermaid** for GitHub documentation. GitHub's Markdown renderer natively supports Mermaid diagrams in fenced code blocks, allowing immediate visualization without external tooling. While Graphviz and PlantUML require additional rendering pipelines or extensions, Mermaid diagrams generated by TextFlow can be viewed directly in pull requests, issues, and README files on GitHub.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →