How the Vision Textualizer Converts Flowchart Images to Structured Text Representations

The Vision Textualizer processes flowchart images through a seven-stage pipeline that encodes visual input, prompts a Vision-Language Model (VLM), and extracts structured diagram code in Mermaid, GraphViz, or PlantUML syntax.

The Vision Textualizer serves as the critical first stage of the TextFlow pipeline, transforming static flowchart images into machine-readable text representations. This component bridges the gap between visual diagrams and downstream processing by leveraging state-of-the-art vision-language models to generate structured code. According to the junyiye/textflow source code, the conversion process follows a rigorous sequence from CLI argument parsing to structured JSON output.

The Seven-Step Conversion Pipeline

The conversion process implemented in src/textualizer.py orchestrates multiple specialized components to transform images into structured text.

1. CLI and Configuration Initialization

The pipeline begins by parsing user arguments and loading system configuration. In src/textualizer.py, the argument parser handles three critical parameters:

parser = argparse.ArgumentParser(
    description="Run the Vision Textualizer program. (Convert flowchart to text Representation)"
)
parser.add_argument("--dataset", default="flowvqa")
parser.add_argument("--textualizer", default="Qwen2-VL-7B")
parser.add_argument("--output_type", default="mermaid")
args = parser.parse_args()

The system loads config.json via load_config() to resolve dataset paths and logging directories. The logger writes timestamped files to config["logging"]["log_dir"], ensuring reproducible execution traces.

2. ModelWrapper Selection

The ModelWrapper class in src/models/model_loader.py determines whether to use an API-based or locally hosted VLM based on the model name:

self.is_api_model = model_name in ["claude-3-5-sonnet", "gpt-4o", "gpt-4o-mini"]
if self.is_api_model:
    self.model = load_api_model(model_name)
else:
    self.model, self.tokenizer = load_local_model(model_name)

This abstraction allows the textualizer to support both cloud-based models (Anthropic Claude, OpenAI GPT-4o) and local inference (Qwen2-VL-7B) through a unified interface.

3. Prompt Engineering

The prompt builder in src/prompts/prompts.py generates format-specific instructions via load_textualizer_prompt(). The function provides concrete examples for each output type:

def load_textualizer_prompt(output_type):
    if output_type == "mermaid":
        return """Generate the Mermaid code for the provided flowchart.

Here is an example:

```mermaid
flowchart TD
    A(["Start"]) --> B[/"Receive 'arr' and 'n'"/]
    ...

```"""

This example-driven approach ensures the VLM produces syntactically valid diagram code in the requested format (Mermaid, GraphViz, or PlantUML).

4. Image Encoding Strategies

The encode_image function in src/utils.py handles provider-specific image preparation requirements:

Model Provider Encoding Method Data Format
Anthropic encode_image_anthropic PNG → base64
OpenAI encode_image_openai Raw bytes → base64
Local Models Image.open() PIL Image object
def encode_image(image_path, model_name=None):
    if model_name == "claude-3-5-sonnet":
        return encode_image_anthropic(image_path)
    elif model_name in ["gpt-4o", "gpt-4o-mini"]:
        return encode_image_openai(image_path)
    else:
        return Image.open(image_path)

This differentiation ensures compatibility with each VLM provider's API requirements.

5. Model Invocation

The wrapped model receives the encoded image and prompt through the generate_response method:

response = model.generate_response(prompt, image_path=image_path)

Internally, ModelWrapper coordinates load_messages to construct the chat payload, then routes to either generate_api_response for cloud models or generate_local_response for on-premise inference. The VLM returns a raw text block containing the diagram wrapped in markdown fences.

6. Structured Text Extraction

The extract_representation function in src/utils.py parses the VLM output to isolate the diagram code:

def extract_representation(string):
    if "```mermaid" in string:
        return extract_mermaid_code(string)
    elif "```dot" in string:
        return extract_graphviz_code(string)
    elif "```plantuml" in string:
        return extract_plantuml_code(string)
    else:
        return string

Specialized regex patterns (e.g., r"```mermaid\s+([\s\S]*?)```") strip the surrounding markdown to return pure diagram syntax.

7. Result Persistence

The final stage writes the mapping of image IDs to diagram code as JSON:

output_dir = os.path.join(config["file_paths"]["output"], dataset, output_type)
output_file = os.path.join(output_dir, f"{textualizer}.json")
os.makedirs(output_dir, exist_ok=True)
with open(output_file, "w") as file:
    json.dump(results, file, indent=4)

Output follows the pattern output/<dataset>/<output_type>/<textualizer>.json, such as output/flowvqa/mermaid/gpt-4o.json.

Practical Usage Examples

Execute the textualizer against the FlowVQA dataset using GPT-4o to generate Mermaid diagrams:

python src/textualizer.py \
    --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

To generate GraphViz syntax using Anthropic's Claude model:

python src/textualizer.py \
    --dataset flowlearn \
    --textualizer claude-3-5-sonnet \
    --output_type graphviz

Summary

  • The Vision Textualizer in src/textualizer.py orchestrates a seven-step pipeline to convert flowchart images into structured text representations.
  • ModelWrapper (src/models/model_loader.py) abstracts API and local VLM execution, supporting Claude-3.5-Sonnet, GPT-4o, and local models like Qwen2-VL-7B.
  • Image encoding adapts to provider requirements: base64 for Anthropic and OpenAI APIs, PIL Image objects for local inference.
  • Prompt engineering in src/prompts/prompts.py uses example-driven instructions to ensure valid Mermaid, GraphViz, or PlantUML output.
  • Response extraction employs regex-based parsing in src/utils.py to isolate diagram code from markdown fences.
  • Output persistence generates JSON files at output/<dataset>/<type>/<model>.json for downstream pipeline stages.

Frequently Asked Questions

What Vision-Language Models does the Vision Textualizer support?

The Vision Textualizer supports both API-based and local models. API models include Claude-3.5-Sonnet, GPT-4o, and GPT-4o-mini, which require base64 image encoding. Local models such as Qwen2-VL-7B run on-premises using standard PIL Image objects without base64 conversion, as implemented in src/models/model_loader.py.

How does the Vision Textualizer handle different diagram syntax formats?

The system uses the load_textualizer_prompt function in src/prompts/prompts.py to inject format-specific examples into the prompt. For Mermaid, it includes flowchart TD examples; for GraphViz, it uses dot syntax examples. The extract_representation function in src/utils.py then detects the corresponding markdown fence (mermaid, dot, or ```plantuml) and extracts the code using regex patterns.

What is the output format of the Vision Textualizer?

The Vision Textualizer produces JSON files mapping image identifiers to extracted diagram code. Files are saved to output/<dataset>/<output_type>/<textualizer>.json, where output_type specifies the diagram language (mermaid, graphviz, or plantuml) and textualizer indicates the VLM used (e.g., gpt-4o.json).

Why does the Vision Textualizer use different image encoding methods?

Different VLM providers require specific input formats. Anthropic's Claude API expects PNG images encoded as base64 strings via encode_image_anthropic, while OpenAI models require raw bytes converted to base64 via encode_image_openai. Local models bypass base64 encoding entirely and receive PIL Image objects directly, optimizing performance for on-device inference as defined in src/utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →