How to Run Baseline VQA Comparison Against the TextFlow Pipeline
Run python src/vqa.py for the baseline and python src/textualizer.py followed by python src/reasoner.py for TextFlow, then evaluate both with python src/evaluation.py to compare accuracy on the flowvqa dataset.
The junyiye/textflow repository provides a complete framework for comparing end-to-end vision models against modular text-based pipelines. This guide walks you through the exact commands and source files needed to execute a baseline VQA comparison against the TextFlow pipeline on the flowvqa benchmark.
Understanding the Baseline and TextFlow Approaches
Baseline VQA (End-to-End)
The baseline approach treats flowchart images as raw visual input. In src/vqa.py, the system loads the test split from data/flowvqa/test.json, constructs a VQA prompt via prompts.load_vqa_prompt(), and queries the vision-language model directly through the ModelWrapper class. This method bypasses any intermediate text representation, producing answers in a single inference step.
TextFlow Pipeline (Image-to-Text-to-Answer)
The TextFlow pipeline decouples visual understanding from reasoning. First, src/textualizer.py converts each flowchart image into a structured text format—Mermaid, GraphViz, or PlantUML—using prompts.load_textualizer_prompt(). Then, src/reasoner.py consumes that textual representation via prompts.load_reasoner_prompt() and applies a language model to derive the final answer. This modular design allows you to swap textualizers or reasoners independently.
Step-by-Step Execution Guide
Step 1 – Run the Baseline VQA
Execute the end-to-end baseline against the flowvqa dataset:
python src/vqa.py --dataset flowvqa --model_name gpt-4o
This command processes data/flowvqa/test.json, invokes the VLM through ModelWrapper, and writes results to output/flowvqa/vqa/gpt-4o.json. The script respects paths defined in config.json and logs progress via src/logger.py.
Step 2 – Execute the TextFlow Textualizer
Generate textual representations from the flowchart images:
python src/textualizer.py --dataset flowvqa --textualizer gpt-4o --output_type mermaid
src/textualizer.py loads the textualizer prompt from src/prompts/prompts.py, queries the vision model, and stores the extracted Mermaid code in output/flowvqa/mermaid/gpt-4o.json. Supported output formats include mermaid, graphviz, and plantuml.
Step 3 – Run the TextFlow Reasoner
Perform question answering over the generated text:
python src/reasoner.py --dataset flowvqa \
--reasoner gpt-4o \
--textualizer gpt-4o \
--input_type mermaid
This command executes src/reasoner.py, which reads the Mermaid output from Step 2, constructs a reasoning prompt via prompts.load_reasoner_prompt(), and runs the language model. Results are saved to output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json.
To enable tool-use for Mermaid parsing, append --tool_use to the reasoner command.
Step 4 – Evaluate Both Results
Compute accuracy for the baseline:
python src/evaluation.py \
--model_name gpt-4o \
--data_path output/flowvqa/vqa/gpt-4o.json
Then evaluate the TextFlow pipeline:
python src/evaluation.py \
--model_name gpt-4o \
--data_path output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json
src/evaluation.py loads the specified result file, generates an evaluation prompt using prompts.load_evaluation_prompt(), obtains three independent judgments from the evaluator LLM, and aggregates them via utils.majority_vote() to produce the final accuracy score.
Complete Command Reference
Execute the full baseline VQA comparison against the TextFlow pipeline in sequence:
# 1. Baseline VQA
python src/vqa.py --dataset flowvqa --model_name gpt-4o
# 2. Textualizer (Mermaid)
python src/textualizer.py --dataset flowvqa \
--textualizer gpt-4o \
--output_type mermaid
# 3. Reasoner (Mermaid, no tool use)
python src/reasoner.py --dataset flowvqa \
--reasoner gpt-4o \
--textualizer gpt-4o \
--input_type mermaid
# 4. Evaluation – Baseline
python src/evaluation.py \
--model_name gpt-4o \
--data_path output/flowvqa/vqa/gpt-4o.json
# 5. Evaluation – TextFlow
python src/evaluation.py \
--model_name gpt-4o \
--data_path output/flowvqa/textflow/mermaid_reasoner_gpt-4o_textualizer_gpt-4o.json
Key Implementation Files
| Component | File | Role |
|---|---|---|
| Baseline VQA script | src/vqa.py |
Executes end-to-end VQA on flowchart images. |
| Vision Textualizer script | src/textualizer.py |
Generates Mermaid, GraphViz, or PlantUML from images. |
| Textual Reasoner script | src/reasoner.py |
Answers questions using the textual representation. |
| Evaluation script | src/evaluation.py |
Judges correctness and computes accuracy via majority vote. |
| Prompt utilities | src/prompts/prompts.py |
Loads VQA, textualizer, reasoner, and evaluation prompts. |
| Configuration | config.json |
Defines file paths and logging settings. |
| Logger helper | src/logger.py |
Centralized logging for all pipeline stages. |
Summary
- Baseline VQA runs directly on images via
src/vqa.py, storing results inoutput/flowvqa/vqa/. - TextFlow pipeline splits the task into
src/textualizer.py(image-to-text) andsrc/reasoner.py(text-to-answer), with outputs inoutput/flowvqa/textflow/. - Unified evaluation uses
src/evaluation.pywith majority voting (utils.majority_vote) to compare accuracy across both approaches. - All scripts share
config.jsonfor paths andsrc/logger.pyfor execution tracking.
Frequently Asked Questions
What is the difference between the baseline VQA and TextFlow pipeline?
The baseline VQA approach feeds flowchart images directly into a vision-language model through src/vqa.py, producing answers in a single inference step. The TextFlow pipeline first converts images to structured text (Mermaid, GraphViz, or PlantUML) using src/textualizer.py, then reasons over that text with src/reasoner.py, separating visual extraction from logical reasoning.
Which output formats does the textualizer support?
According to src/textualizer.py, the system supports three output formats specified via the --output_type argument: Mermaid, GraphViz, and PlantUML. The default and recommended format is Mermaid, which integrates with the optional tool-use functionality in the reasoner stage.
How does the evaluation script determine correctness?
src/evaluation.py judges answers by loading the specified result file, constructing an evaluation prompt via prompts.load_evaluation_prompt(), and querying the evaluator LLM for three independent judgments. It then aggregates these judgments using utils.majority_vote() to produce a final binary correctness label and calculates overall accuracy across the dataset.
Can I use different models for the textualizer and reasoner?
Yes. The TextFlow pipeline decouples these stages, allowing you to specify different models via the --textualizer and --reasoner arguments. For example, you can generate Mermaid diagrams with gpt-4o and reason over them with claude-3-opus-20240229, enabling modular experimentation with vision and language model combinations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →