How to Visualize Intermediate Text Representations for Debugging in TextFlow

You can visualize intermediate text representations in TextFlow by inspecting the JSON files in the output/ directory, enabling the built-in logger in src/textualizer.py, or rendering the extracted Mermaid, Graphviz, or PlantUML code with external diagram viewers.

TextFlow, an open-source pipeline that converts flowchart images into structured text, generates intermediate representations during its Vision Textualizer stage. Understanding how to access and visualize these artifacts is essential for debugging the conversion process before the Textual Reasoner consumes the output.

Where TextFlow Stores Intermediate Representations

The intermediate representation is produced in src/textualizer.py by calling a Vision Language Model (VLM), then extracted using utility functions from src/utils.py. After each image is processed, the representation is serialized to a JSON file under the output/ directory (see src/textualizer.py lines 76‑80).

Each JSON file follows the structure where keys are image IDs and values contain the raw textual code (Mermaid, Graphviz DOT, or PlantUML syntax). This same JSON is later loaded by src/reasoner.py (lines 81‑88), making it the critical handoff point between pipeline stages.

Three Methods to Visualize Intermediate Text Representations for Debugging

Inspect the JSON Output Directly

The simplest way to visualize intermediate text representations is to parse the generated JSON file. Each entry contains the complete diagram code extracted from the flowchart image.

from pathlib import Path
import json

# Path to the generated JSON (adjust dataset, type, and model as needed)

out_path = Path("output/flowvqa/mermaid/Qwen2-VL-7B.json")
data = json.load(out_path.open())

# Display the representation for a specific sample

sample_id = "0"
print(f"Extracted code for sample {sample_id}:\n{data[sample_id]}")

Relevant source: src/textualizer.py – generation loop (lines 66‑74) and JSON dump (lines 76‑80).

Enable Logging for Real-Time Debugging

TextFlow provides a setup_logger function that writes extracted representations to a log file during execution. This is useful for monitoring the pipeline without manually parsing JSON files.

The logger is initialized in src/textualizer.py (lines 52‑55) and records the intermediate representation immediately after extraction. Check the log file specified in your configuration to see the raw text output alongside timestamps and image IDs.

Render Diagrams with External Tools

To verify that the extracted text accurately represents the original flowchart, render the intermediate representation using format-specific tools.

Mermaid Diagrams

Save the extracted code to an HTML file with the Mermaid.js CDN:

<!DOCTYPE html>
<html>
<head>
  <script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
  <script>mermaid.initialize({startOnLoad:true});</script>
</head>
<body>
  <div class="mermaid">
graph TD;
    A[Start] --> B{Decision};
    B -->|Yes| C[Action 1];
    B -->|No| D[Action 2];
  </div>
</body>
</html>

Replace the graph content with your extracted Mermaid code from the JSON file.

Graphviz (DOT) Diagrams

For Graphviz output, use the command-line tool to generate an image:


# Assuming DOT code is saved in diagram.dot

dot -Tpng diagram.dot -o diagram.png

Relevant source: src/utils.py – extract_graphviz_code (lines 60‑70).

PlantUML Diagrams

For PlantUML syntax, encode the text and query the public PlantUML server:


# Encode the UML code

encoded=$(python -c "import urllib.parse, sys; print(urllib.parse.quote(sys.stdin.read()))" < uml_code.txt)

# Fetch SVG from PlantUML server

curl -s "https://www.plantuml.com/plantuml/svg/${encoded}" > diagram.svg

Relevant source: src/utils.py – extract_plantuml_code (lines 73‑83).

How the Reasoner Consumes These Representations

Understanding the data flow helps debug mismatches between what you visualize and what the pipeline processes. The src/reasoner.py module loads the JSON file (lines 81‑88) produced by the textualizer and passes the extracted strings to the LLM reasoner.

If the reasoner fails to parse the diagram logic, compare the raw string in the JSON file against what the reasoner receives. Because both stages reference the same output/<dataset>/<type>/<model>.json file, any discrepancy indicates a file I/O or encoding issue rather than a model error.

Summary

  • TextFlow generates intermediate text representations in Mermaid, Graphviz, or PlantUML formats during the Vision Textualizer stage.
  • These representations are serialized to JSON files in the output/ directory by src/textualizer.py (lines 76‑80).
  • You can visualize intermediate text representations for debugging by inspecting the JSON directly, enabling the logger (lines 52‑55), or rendering the code with external tools like Mermaid Live Editor, Graphviz CLI, or PlantUML servers.
  • The src/reasoner.py module consumes the same JSON files (lines 81‑88), allowing you to verify data consistency between pipeline stages.

Frequently Asked Questions

What file formats does TextFlow use for intermediate representations?

TextFlow supports three textual diagram formats: Mermaid for flowchart syntax, Graphviz (DOT) for graph descriptions, and PlantUML for UML diagrams. You specify the desired output type using the --output_type parameter when running src/textualizer.py.

Where are the intermediate representations saved in the TextFlow pipeline?

The intermediate representations are saved as JSON files in the output/<dataset>/<type>/<model>.json path structure. This serialization occurs in src/textualizer.py at lines 76‑80, where the script writes the extracted diagram code mapped to image IDs.

How can I view Mermaid diagrams generated by TextFlow?

You can view Mermaid diagrams by extracting the code from the JSON output and rendering it with the Mermaid Live Editor online, or by embedding the code in an HTML file using the Mermaid.js CDN. Alternatively, save the code to a .mmd file and use the Mermaid CLI to generate PNG or SVG images.

Why is my intermediate representation not being processed by the reasoner?

If the reasoner fails to process the representation, verify that the JSON file exists in the expected output/ subdirectory and that the image ID keys match between the textualizer output and reasoner input. Check src/reasoner.py lines 81‑88 to confirm the file loading logic, and ensure no encoding errors corrupted the diagram syntax during the extraction process in src/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →