Datasets Supported by TextFlow and How to Prepare Custom Flowchart Data

TextFlow supports three built-in benchmark datasets—FlowVQA, FlowVQA Bottom-Top, and FlowLearn—and accepts custom datasets by placing images and a structured JSON file in a new folder registered in config.json.

TextFlow is an open-source framework for visual question answering on flowchart images. Understanding the datasets supported by TextFlow is essential for running experiments and extending the pipeline to your own flowchart collections. This guide covers the built-in benchmarks, their file locations, and the exact schema required to prepare custom data.

Built-in Datasets Supported by TextFlow

TextFlow ships with three pre-configured benchmark datasets located in the data/ directory. Each dataset follows a consistent structure containing a test.json file and an images/ subdirectory.

FlowVQA

FlowVQA is the primary benchmark containing full-size flowchart images and question-answer pairs. It is the default dataset used by the training and evaluation scripts.

  • Repository path: data/flowvqa/
  • Data layout: test.json contains the QA annotations; images are stored in images/ as PNG files
  • CLI usage: --dataset flowvqa

FlowVQA Bottom-Top

FlowVQA Bottom-Top is a variant of FlowVQA where each flowchart image is split into bottom and top halves. This dataset is used for ablation studies to evaluate model performance on partial visual information.

  • Repository path: data/flowvqa_bottom_top/
  • Data layout: Same JSON structure as FlowVQA with corresponding split images
  • CLI usage: --dataset flowvqa_bottom_top

FlowLearn

FlowLearn is a smaller benchmark focused on learning the visual-text mapping for flowcharts. It uses JPEG images and is ideal for rapid prototyping and debugging the textualizer component.

  • Repository path: data/flowlearn/
  • Data layout: test.json with JPEG images in images/
  • CLI usage: --dataset flowlearn

Dataset Configuration in config.json

The list of datasets supported by TextFlow is centrally managed in config.json at the repository root. The "file_paths" key maps dataset names to their directories:

{
  "file_paths": {
    "flowvqa": "data/flowvqa",
    "flowvqa_bottom_top": "data/flowvqa_bottom_top",
    "flowlearn": "data/flowlearn",
    "output": "output"
  }
}

This configuration is loaded by the CLI scripts (src/vqa.py, src/textualizer.py, src/reasoner.py) to resolve dataset locations. The --dataset argument defaults to these keys.

How to Prepare Your Own Flowchart Data

To extend TextFlow beyond the built-in benchmarks, you must create a dataset folder that mirrors the existing structure and follows the exact JSON schema expected by the data loader.

Required Directory Structure

Create a new folder under data/ with the following layout:

data/my_flowcharts/
├── test.json
└── images/
    ├── chart_001.png
    ├── chart_002.png
    └── ...

JSON Schema for test.json

The test.json file must contain a dictionary where each key is a unique identifier matching the image filename (without extension). The value must follow this schema:

{
  "<unique_id>": {
    "key": "<same_as_unique_id>",
    "title": "<human_readable_title>",
    "text": "<optional_raw_paragraph>",
    "category1": "<high_level_category>",
    "category2": "<sub_category>",
    "tags": "[...]",
    "summary": "<optional_markdown_summary>",
    "mermaid": "<Mermaid_flowchart_string>",
    "graphviz": "<GraphViz_DOT_string_optional>",
    "plantuml": "<PlantUML_string_optional>",
    "qa": {
      "<qa_id>": {
        "Q": "<question_text>",
        "A1": "<gold_answer_1>",
        "A2": "<gold_answer_2_optional>",
        "A3": "<gold_answer_3_optional>",
        "type": "<question_type>"
      }
    }
  }
}

Critical fields:

  • key: Must match the dictionary key and the image filename.
  • mermaid: Required for the textualizer to generate flowchart text (unless using graphviz or plantuml with matching --output_type).
  • qa: Must contain at least Q and A1 for evaluation.

Image File Requirements

  • Format: Use .png for standard datasets (matching FlowVQA) or .jpeg for FlowLearn-style data.
  • Naming: Image filenames must exactly match the key field in test.json (e.g., my_chart.png for key "my_chart").
  • Location: Place all images in the images/ subdirectory of your dataset folder.

Registering Your Custom Dataset

After creating the folder and JSON file, register the dataset in config.json:

{
  "file_paths": {
    "flowvqa": "data/flowvqa",
    "flowvqa_bottom_top": "data/flowvqa_bottom_top",
    "flowlearn": "data/flowlearn",
    "my_flowcharts": "data/my_flowcharts",
    "output": "output"
  }
}

You can now use --dataset my_flowcharts in any TextFlow CLI command.

Running the TextFlow Pipeline on Custom Data

Once your dataset is prepared and registered, you can execute the three-stage pipeline using the scripts in src/.

1. Vision Question Answering (VQA)

Run the vision-only baseline using src/vqa.py:

python src/vqa.py \
  --dataset my_flowcharts \
  --model_name gpt-4o

This loads images from data/my_flowcharts/images/ and annotations from test.json, then writes results to output/my_flowcharts/vqa/.

2. Vision Textualizer – Generate Mermaid Representation

Convert images to textual flowchart representations using src/textualizer.py:

python src/textualizer.py \
  --dataset my_flowcharts \
  --textualizer Qwen2-VL-7B \
  --output_type mermaid

The script uses the VLM to generate Mermaid syntax, extracts it using utils.extract_representation, and saves the output to output/my_flowcharts/mermaid/my_flowcharts.json.

3. Textual Reasoner – Answer Questions Using Generated Text

Perform reasoning over the textual representation using src/reasoner.py:

python src/reasoner.py \
  --dataset my_flowcharts \
  --reasoner Llama-3.1-8B \
  --textualizer Qwen2-VL-7B \
  --input_type mermaid

The reasoner loads prompts from prompts.load_reasoner_prompt, processes the Mermaid text generated in the previous step, and writes answers to output/my_flowcharts/textflow/mermaid_reasoner_Llama-3.1-8B_textualizer_Qwen2-VL-7B.json.

Summary

  • TextFlow supports three built-in datasets: FlowVQA, FlowVQA Bottom-Top, and FlowLearn, located in data/flowvqa/, data/flowvqa_bottom_top/, and data/flowlearn/ respectively.
  • Dataset registration occurs in config.json under the "file_paths" key, which maps dataset names to directory paths.
  • Custom datasets require a folder containing test.json (following the specific schema with key, mermaid, and qa fields) and an images/ subdirectory with PNG or JPEG files matching the JSON keys.
  • The pipeline consists of three stages—src/vqa.py for vision QA, src/textualizer.py for generating Mermaid/Graphviz text, and src/reasoner.py for textual reasoning—each accepting the --dataset argument to target custom data.

Frequently Asked Questions

What file format should I use for flowchart images in a custom TextFlow dataset?

TextFlow accepts PNG files for standard datasets like FlowVQA and JPEG files for FlowLearn-style data. When preparing your own dataset, use PNG unless you are specifically matching the FlowLearn benchmark format. The image filename (without extension) must exactly match the key field in your test.json file.

Do I need to provide all three text representations (Mermaid, Graphviz, and PlantUML) in my custom dataset?

No, you only need to provide at least one text representation. The mermaid field is the most commonly used and is required if you plan to use the default --output_type mermaid with src/textualizer.py. If you provide graphviz or plantuml instead, ensure you set the matching --output_type parameter when running the textualizer script.

How does TextFlow handle question types in the QA section?

The type field inside each QA object guides the evaluation logic. Common values include "topological" for questions about node or edge counting, "fact_retrieval" for extracting specific information from the flowchart, and "applied_scenario" for reasoning questions. While the field is required for proper evaluation, you can define custom types as long as your evaluation script handles them accordingly.

Can I use a custom dataset name with spaces or special characters?

It is recommended to use lowercase alphanumeric characters and underscores for dataset names (e.g., my_flowcharts or custom_v1). The dataset name is used as a dictionary key in config.json and as a folder name in the file system. Special characters or spaces may cause issues with the CLI argument parser or file path resolution in src/vqa.py, src/textualizer.py, and src/reasoner.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →