# How the Vision Textualizer Converts Flowchart Images to Structured Text Representations

> Learn how the Vision Textualizer transforms flowchart images into structured text like Mermaid or GraphViz. Discover its seven-stage pipeline and VLM integration for diagram code.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**The Vision Textualizer processes flowchart images through a seven-stage pipeline that encodes visual input, prompts a Vision-Language Model (VLM), and extracts structured diagram code in Mermaid, GraphViz, or PlantUML syntax.**

The Vision Textualizer serves as the critical first stage of the TextFlow pipeline, transforming static flowchart images into machine-readable text representations. This component bridges the gap between visual diagrams and downstream processing by leveraging state-of-the-art vision-language models to generate structured code. According to the junyiye/textflow source code, the conversion process follows a rigorous sequence from CLI argument parsing to structured JSON output.

## The Seven-Step Conversion Pipeline

The conversion process implemented in [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) orchestrates multiple specialized components to transform images into structured text.

### 1. CLI and Configuration Initialization

The pipeline begins by parsing user arguments and loading system configuration. In [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), the argument parser handles three critical parameters:

```python
parser = argparse.ArgumentParser(
    description="Run the Vision Textualizer program. (Convert flowchart to text Representation)"
)
parser.add_argument("--dataset", default="flowvqa")
parser.add_argument("--textualizer", default="Qwen2-VL-7B")
parser.add_argument("--output_type", default="mermaid")
args = parser.parse_args()

```

The system loads [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) via `load_config()` to resolve dataset paths and logging directories. The logger writes timestamped files to `config["logging"]["log_dir"]`, ensuring reproducible execution traces.

### 2. ModelWrapper Selection

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py) determines whether to use an API-based or locally hosted VLM based on the model name:

```python
self.is_api_model = model_name in ["claude-3-5-sonnet", "gpt-4o", "gpt-4o-mini"]
if self.is_api_model:
    self.model = load_api_model(model_name)
else:
    self.model, self.tokenizer = load_local_model(model_name)

```

This abstraction allows the textualizer to support both cloud-based models (Anthropic Claude, OpenAI GPT-4o) and local inference (Qwen2-VL-7B) through a unified interface.

### 3. Prompt Engineering

The prompt builder in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) generates format-specific instructions via `load_textualizer_prompt()`. The function provides concrete examples for each output type:

```python
def load_textualizer_prompt(output_type):
    if output_type == "mermaid":
        return """Generate the Mermaid code for the provided flowchart.

Here is an example:

```mermaid
flowchart TD
    A(["Start"]) --> B[/"Receive 'arr' and 'n'"/]
    ...

```"""

```

This example-driven approach ensures the VLM produces syntactically valid diagram code in the requested format (Mermaid, GraphViz, or PlantUML).

### 4. Image Encoding Strategies

The `encode_image` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) handles provider-specific image preparation requirements:

| Model Provider | Encoding Method | Data Format |
| --- | --- | --- |
| **Anthropic** | `encode_image_anthropic` | PNG → base64 |
| **OpenAI** | `encode_image_openai` | Raw bytes → base64 |
| **Local Models** | `Image.open()` | PIL Image object |

```python
def encode_image(image_path, model_name=None):
    if model_name == "claude-3-5-sonnet":
        return encode_image_anthropic(image_path)
    elif model_name in ["gpt-4o", "gpt-4o-mini"]:
        return encode_image_openai(image_path)
    else:
        return Image.open(image_path)

```

This differentiation ensures compatibility with each VLM provider's API requirements.

### 5. Model Invocation

The wrapped model receives the encoded image and prompt through the `generate_response` method:

```python
response = model.generate_response(prompt, image_path=image_path)

```

Internally, `ModelWrapper` coordinates `load_messages` to construct the chat payload, then routes to either `generate_api_response` for cloud models or `generate_local_response` for on-premise inference. The VLM returns a raw text block containing the diagram wrapped in markdown fences.

### 6. Structured Text Extraction

The `extract_representation` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) parses the VLM output to isolate the diagram code:

```python
def extract_representation(string):
    if "```mermaid" in string:
        return extract_mermaid_code(string)
    elif "```dot" in string:
        return extract_graphviz_code(string)
    elif "```plantuml" in string:
        return extract_plantuml_code(string)
    else:
        return string

```

Specialized regex patterns (e.g., `r"```mermaid\s+([\s\S]*?)```"`) strip the surrounding markdown to return pure diagram syntax.

### 7. Result Persistence

The final stage writes the mapping of image IDs to diagram code as JSON:

```python
output_dir = os.path.join(config["file_paths"]["output"], dataset, output_type)
output_file = os.path.join(output_dir, f"{textualizer}.json")
os.makedirs(output_dir, exist_ok=True)
with open(output_file, "w") as file:
    json.dump(results, file, indent=4)

```

Output follows the pattern `output/<dataset>/<output_type>/<textualizer>.json`, such as `output/flowvqa/mermaid/gpt-4o.json`.

## Practical Usage Examples

Execute the textualizer against the FlowVQA dataset using GPT-4o to generate Mermaid diagrams:

```bash
python src/textualizer.py \
    --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

```

To generate GraphViz syntax using Anthropic's Claude model:

```bash
python src/textualizer.py \
    --dataset flowlearn \
    --textualizer claude-3-5-sonnet \
    --output_type graphviz

```

## Summary

- **The Vision Textualizer** in `src/textualizer.py` orchestrates a seven-step pipeline to convert flowchart images into structured text representations.
- **ModelWrapper** (`src/models/model_loader.py`) abstracts API and local VLM execution, supporting Claude-3.5-Sonnet, GPT-4o, and local models like Qwen2-VL-7B.
- **Image encoding** adapts to provider requirements: base64 for Anthropic and OpenAI APIs, PIL Image objects for local inference.
- **Prompt engineering** in `src/prompts/prompts.py` uses example-driven instructions to ensure valid Mermaid, GraphViz, or PlantUML output.
- **Response extraction** employs regex-based parsing in `src/utils.py` to isolate diagram code from markdown fences.
- **Output persistence** generates JSON files at `output/<dataset>/<type>/<model>.json` for downstream pipeline stages.

## Frequently Asked Questions

### What Vision-Language Models does the Vision Textualizer support?

The Vision Textualizer supports both API-based and local models. API models include **Claude-3.5-Sonnet**, **GPT-4o**, and **GPT-4o-mini**, which require base64 image encoding. Local models such as **Qwen2-VL-7B** run on-premises using standard PIL Image objects without base64 conversion, as implemented in `src/models/model_loader.py`.

### How does the Vision Textualizer handle different diagram syntax formats?

The system uses the `load_textualizer_prompt` function in `src/prompts/prompts.py` to inject format-specific examples into the prompt. For Mermaid, it includes `flowchart TD` examples; for GraphViz, it uses `dot` syntax examples. The `extract_representation` function in `src/utils.py` then detects the corresponding markdown fence (```mermaid, ```dot, or ```plantuml) and extracts the code using regex patterns.

### What is the output format of the Vision Textualizer?

The Vision Textualizer produces JSON files mapping image identifiers to extracted diagram code. Files are saved to `output/<dataset>/<output_type>/<textualizer>.json`, where `output_type` specifies the diagram language (mermaid, graphviz, or plantuml) and `textualizer` indicates the VLM used (e.g., [`gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/gpt-4o.json)).

### Why does the Vision Textualizer use different image encoding methods?

Different VLM providers require specific input formats. Anthropic's Claude API expects PNG images encoded as base64 strings via `encode_image_anthropic`, while OpenAI models require raw bytes converted to base64 via `encode_image_openai`. Local models bypass base64 encoding entirely and receive PIL Image objects directly, optimizing performance for on-device inference as defined in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py).