# How to Run Baseline VQA Comparison Against the TextFlow Pipeline

> Compare VQA models with the TextFlow pipeline. Learn how to run the baseline VQA and TextFlow comparison for accurate results using junyiye/textflow.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Run `python src/vqa.py` for the baseline and `python src/textualizer.py` followed by `python src/reasoner.py` for TextFlow, then evaluate both with `python src/evaluation.py` to compare accuracy on the flowvqa dataset.**

The `junyiye/textflow` repository provides a complete framework for comparing end-to-end vision models against modular text-based pipelines. This guide walks you through the exact commands and source files needed to execute a baseline VQA comparison against the TextFlow pipeline on the flowvqa benchmark.

## Understanding the Baseline and TextFlow Approaches

### Baseline VQA (End-to-End)

The baseline approach treats flowchart images as raw visual input. In [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), the system loads the test split from [`data/flowvqa/test.json`](https://github.com/junyiye/textflow/blob/main/data/flowvqa/test.json), constructs a VQA prompt via `prompts.load_vqa_prompt()`, and queries the vision-language model directly through the `ModelWrapper` class. This method bypasses any intermediate text representation, producing answers in a single inference step.

### TextFlow Pipeline (Image-to-Text-to-Answer)

The TextFlow pipeline decouples visual understanding from reasoning. First, [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) converts each flowchart image into a structured text format—Mermaid, GraphViz, or PlantUML—using `prompts.load_textualizer_prompt()`. Then, [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) consumes that textual representation via `prompts.load_reasoner_prompt()` and applies a language model to derive the final answer. This modular design allows you to swap textualizers or reasoners independently.

## Step-by-Step Execution Guide

### Step 1 – Run the Baseline VQA

Execute the end-to-end baseline against the flowvqa dataset:

```bash
python src/vqa.py --dataset flowvqa --model_name gpt-4o

```

This command processes [`data/flowvqa/test.json`](https://github.com/junyiye/textflow/blob/main/data/flowvqa/test.json), invokes the VLM through `ModelWrapper`, and writes results to [`output/flowvqa/vqa/gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/output/flowvqa/vqa/gpt-4o.json). The script respects paths defined in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) and logs progress via [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py).

### Step 2 – Execute the TextFlow Textualizer

Generate textual representations from the flowchart images:

```bash
python src/textualizer.py --dataset flowvqa --textualizer gpt-4o --output_type mermaid

```

[`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) loads the textualizer prompt from [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py), queries the vision model, and stores the extracted Mermaid code in [`output/flowvqa/mermaid/gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/output/flowvqa/mermaid/gpt-4o.json). Supported output formats include `mermaid`, `graphviz`, and `plantuml`.

### Step 3 – Run the TextFlow Reasoner

Perform question answering over the generated text:

```bash
python src/reasoner.py --dataset flowvqa \
    --reasoner gpt-4o \
    --textualizer gpt-4o \
    --input_type mermaid

```

This command executes [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), which reads the Mermaid output from Step 2, constructs a reasoning prompt via `prompts.load_reasoner_prompt()`, and runs the language model. Results are saved to [`output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json`](https://github.com/junyiye/textflow/blob/main/output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json).

To enable tool-use for Mermaid parsing, append `--tool_use` to the reasoner command.

### Step 4 – Evaluate Both Results

Compute accuracy for the baseline:

```bash
python src/evaluation.py \
    --model_name gpt-4o \
    --data_path output/flowvqa/vqa/gpt-4o.json

```

Then evaluate the TextFlow pipeline:

```bash
python src/evaluation.py \
    --model_name gpt-4o \
    --data_path output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json

```

[`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) loads the specified result file, generates an evaluation prompt using `prompts.load_evaluation_prompt()`, obtains three independent judgments from the evaluator LLM, and aggregates them via `utils.majority_vote()` to produce the final accuracy score.

## Complete Command Reference

Execute the full baseline VQA comparison against the TextFlow pipeline in sequence:

```bash

# 1. Baseline VQA

python src/vqa.py --dataset flowvqa --model_name gpt-4o

# 2. Textualizer (Mermaid)

python src/textualizer.py --dataset flowvqa \
    --textualizer gpt-4o \
    --output_type mermaid

# 3. Reasoner (Mermaid, no tool use)

python src/reasoner.py --dataset flowvqa \
    --reasoner gpt-4o \
    --textualizer gpt-4o \
    --input_type mermaid

# 4. Evaluation – Baseline

python src/evaluation.py \
    --model_name gpt-4o \
    --data_path output/flowvqa/vqa/gpt-4o.json

# 5. Evaluation – TextFlow

python src/evaluation.py \
    --model_name gpt-4o \
    --data_path output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json

```

## Key Implementation Files

| Component | File | Role |
|-----------|------|------|
| Baseline VQA script | [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) | Executes end-to-end VQA on flowchart images. |
| Vision Textualizer script | [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) | Generates Mermaid, GraphViz, or PlantUML from images. |
| Textual Reasoner script | [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) | Answers questions using the textual representation. |
| Evaluation script | [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) | Judges correctness and computes accuracy via majority vote. |
| Prompt utilities | [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) | Loads VQA, textualizer, reasoner, and evaluation prompts. |
| Configuration | [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) | Defines file paths and logging settings. |
| Logger helper | [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py) | Centralized logging for all pipeline stages. |

## Summary

- **Baseline VQA** runs directly on images via [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), storing results in `output/flowvqa/vqa/`.
- **TextFlow pipeline** splits the task into [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) (image-to-text) and [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) (text-to-answer), with outputs in `output/flowvqa/textflow/`.
- **Unified evaluation** uses [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) with majority voting (`utils.majority_vote`) to compare accuracy across both approaches.
- All scripts share [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) for paths and [`src/logger.py`](https://github.com/junyiye/textflow/blob/main/src/logger.py) for execution tracking.

## Frequently Asked Questions

### What is the difference between the baseline VQA and TextFlow pipeline?

The baseline VQA approach feeds flowchart images directly into a vision-language model through [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), producing answers in a single inference step. The TextFlow pipeline first converts images to structured text (Mermaid, GraphViz, or PlantUML) using [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), then reasons over that text with [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), separating visual extraction from logical reasoning.

### Which output formats does the textualizer support?

According to [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), the system supports three output formats specified via the `--output_type` argument: **Mermaid**, **GraphViz**, and **PlantUML**. The default and recommended format is Mermaid, which integrates with the optional tool-use functionality in the reasoner stage.

### How does the evaluation script determine correctness?

[`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) judges answers by loading the specified result file, constructing an evaluation prompt via `prompts.load_evaluation_prompt()`, and querying the evaluator LLM for three independent judgments. It then aggregates these judgments using `utils.majority_vote()` to produce a final binary correctness label and calculates overall accuracy across the dataset.

### Can I use different models for the textualizer and reasoner?

Yes. The TextFlow pipeline decouples these stages, allowing you to specify different models via the `--textualizer` and `--reasoner` arguments. For example, you can generate Mermaid diagrams with `gpt-4o` and reason over them with `claude-3-opus-20240229`, enabling modular experimentation with vision and language model combinations.