# Datasets Supported by TextFlow and How to Prepare Custom Flowchart Data

> Discover TextFlow datasets and learn to prepare custom flowchart data. Support for FlowVQA, FlowVQA Bottom-Top, and FlowLearn, plus easy custom dataset integration.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**TextFlow supports three built-in benchmark datasets—FlowVQA, FlowVQA Bottom-Top, and FlowLearn—and accepts custom datasets by placing images and a structured JSON file in a new folder registered in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json).**

TextFlow is an open-source framework for visual question answering on flowchart images. Understanding the datasets supported by TextFlow is essential for running experiments and extending the pipeline to your own flowchart collections. This guide covers the built-in benchmarks, their file locations, and the exact schema required to prepare custom data.

## Built-in Datasets Supported by TextFlow

TextFlow ships with three pre-configured benchmark datasets located in the `data/` directory. Each dataset follows a consistent structure containing a [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) file and an `images/` subdirectory.

### FlowVQA

**FlowVQA** is the primary benchmark containing full-size flowchart images and question-answer pairs. It is the default dataset used by the training and evaluation scripts.

- **Repository path:** `data/flowvqa/`
- **Data layout:** [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) contains the QA annotations; images are stored in `images/` as PNG files
- **CLI usage:** `--dataset flowvqa`

### FlowVQA Bottom-Top

**FlowVQA Bottom-Top** is a variant of FlowVQA where each flowchart image is split into *bottom* and *top* halves. This dataset is used for ablation studies to evaluate model performance on partial visual information.

- **Repository path:** `data/flowvqa_bottom_top/`
- **Data layout:** Same JSON structure as FlowVQA with corresponding split images
- **CLI usage:** `--dataset flowvqa_bottom_top`

### FlowLearn

**FlowLearn** is a smaller benchmark focused on learning the visual-text mapping for flowcharts. It uses JPEG images and is ideal for rapid prototyping and debugging the textualizer component.

- **Repository path:** `data/flowlearn/`
- **Data layout:** [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) with JPEG images in `images/`
- **CLI usage:** `--dataset flowlearn`

## Dataset Configuration in config.json

The list of datasets supported by TextFlow is centrally managed in **[`config.json`](https://github.com/junyiye/textflow/blob/main/config.json)** at the repository root. The `"file_paths"` key maps dataset names to their directories:

```json
{
  "file_paths": {
    "flowvqa": "data/flowvqa",
    "flowvqa_bottom_top": "data/flowvqa_bottom_top",
    "flowlearn": "data/flowlearn",
    "output": "output"
  }
}

```

This configuration is loaded by the CLI scripts ([`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py)) to resolve dataset locations. The `--dataset` argument defaults to these keys.

## How to Prepare Your Own Flowchart Data

To extend TextFlow beyond the built-in benchmarks, you must create a dataset folder that mirrors the existing structure and follows the exact JSON schema expected by the data loader.

### Required Directory Structure

Create a new folder under `data/` with the following layout:

```bash
data/my_flowcharts/
├── test.json
└── images/
    ├── chart_001.png
    ├── chart_002.png
    └── ...

```

### JSON Schema for test.json

The [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) file must contain a **dictionary** where each key is a unique identifier matching the image filename (without extension). The value must follow this schema:

```json
{
  "<unique_id>": {
    "key": "<same_as_unique_id>",
    "title": "<human_readable_title>",
    "text": "<optional_raw_paragraph>",
    "category1": "<high_level_category>",
    "category2": "<sub_category>",
    "tags": "[...]",
    "summary": "<optional_markdown_summary>",
    "mermaid": "<Mermaid_flowchart_string>",
    "graphviz": "<GraphViz_DOT_string_optional>",
    "plantuml": "<PlantUML_string_optional>",
    "qa": {
      "<qa_id>": {
        "Q": "<question_text>",
        "A1": "<gold_answer_1>",
        "A2": "<gold_answer_2_optional>",
        "A3": "<gold_answer_3_optional>",
        "type": "<question_type>"
      }
    }
  }
}

```

**Critical fields:**
- **key**: Must match the dictionary key and the image filename.
- **mermaid**: Required for the textualizer to generate flowchart text (unless using `graphviz` or `plantuml` with matching `--output_type`).
- **qa**: Must contain at least `Q` and `A1` for evaluation.

### Image File Requirements

- **Format**: Use `.png` for standard datasets (matching FlowVQA) or `.jpeg` for FlowLearn-style data.
- **Naming**: Image filenames must exactly match the `key` field in [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) (e.g., `my_chart.png` for key `"my_chart"`).
- **Location**: Place all images in the `images/` subdirectory of your dataset folder.

### Registering Your Custom Dataset

After creating the folder and JSON file, register the dataset in **[`config.json`](https://github.com/junyiye/textflow/blob/main/config.json)**:

```json
{
  "file_paths": {
    "flowvqa": "data/flowvqa",
    "flowvqa_bottom_top": "data/flowvqa_bottom_top",
    "flowlearn": "data/flowlearn",
    "my_flowcharts": "data/my_flowcharts",
    "output": "output"
  }
}

```

You can now use `--dataset my_flowcharts` in any TextFlow CLI command.

## Running the TextFlow Pipeline on Custom Data

Once your dataset is prepared and registered, you can execute the three-stage pipeline using the scripts in `src/`.

### 1. Vision Question Answering (VQA)

Run the vision-only baseline using [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py):

```bash
python src/vqa.py \
  --dataset my_flowcharts \
  --model_name gpt-4o

```

This loads images from `data/my_flowcharts/images/` and annotations from [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json), then writes results to `output/my_flowcharts/vqa/`.

### 2. Vision Textualizer – Generate Mermaid Representation

Convert images to textual flowchart representations using [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py):

```bash
python src/textualizer.py \
  --dataset my_flowcharts \
  --textualizer Qwen2-VL-7B \
  --output_type mermaid

```

The script uses the VLM to generate Mermaid syntax, extracts it using `utils.extract_representation`, and saves the output to [`output/my_flowcharts/mermaid/my_flowcharts.json`](https://github.com/junyiye/textflow/blob/main/output/my_flowcharts/mermaid/my_flowcharts.json).

### 3. Textual Reasoner – Answer Questions Using Generated Text

Perform reasoning over the textual representation using [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py):

```bash
python src/reasoner.py \
  --dataset my_flowcharts \
  --reasoner Llama-3.1-8B \
  --textualizer Qwen2-VL-7B \
  --input_type mermaid

```

The reasoner loads prompts from `prompts.load_reasoner_prompt`, processes the Mermaid text generated in the previous step, and writes answers to [`output/my_flowcharts/textflow/mermaid_reasoner_Llama-3.1-8B_textualizer_Qwen2-VL-7B.json`](https://github.com/junyiye/textflow/blob/main/output/my_flowcharts/textflow/mermaid_reasoner_Llama-3.1-8B_textualizer_Qwen2-VL-7B.json).

## Summary

- **TextFlow supports three built-in datasets**: FlowVQA, FlowVQA Bottom-Top, and FlowLearn, located in `data/flowvqa/`, `data/flowvqa_bottom_top/`, and `data/flowlearn/` respectively.
- **Dataset registration** occurs in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) under the `"file_paths"` key, which maps dataset names to directory paths.
- **Custom datasets** require a folder containing [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) (following the specific schema with `key`, `mermaid`, and `qa` fields) and an `images/` subdirectory with PNG or JPEG files matching the JSON keys.
- **The pipeline** consists of three stages—[`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py) for vision QA, [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py) for generating Mermaid/Graphviz text, and [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) for textual reasoning—each accepting the `--dataset` argument to target custom data.

## Frequently Asked Questions

### What file format should I use for flowchart images in a custom TextFlow dataset?

TextFlow accepts **PNG** files for standard datasets like FlowVQA and **JPEG** files for FlowLearn-style data. When preparing your own dataset, use PNG unless you are specifically matching the FlowLearn benchmark format. The image filename (without extension) must exactly match the `key` field in your [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) file.

### Do I need to provide all three text representations (Mermaid, Graphviz, and PlantUML) in my custom dataset?

No, you only need to provide **at least one** text representation. The `mermaid` field is the most commonly used and is required if you plan to use the default `--output_type mermaid` with [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py). If you provide `graphviz` or `plantuml` instead, ensure you set the matching `--output_type` parameter when running the textualizer script.

### How does TextFlow handle question types in the QA section?

The `type` field inside each QA object guides the evaluation logic. Common values include `"topological"` for questions about node or edge counting, `"fact_retrieval"` for extracting specific information from the flowchart, and `"applied_scenario"` for reasoning questions. While the field is required for proper evaluation, you can define custom types as long as your evaluation script handles them accordingly.

### Can I use a custom dataset name with spaces or special characters?

It is recommended to use **lowercase alphanumeric characters and underscores** for dataset names (e.g., `my_flowcharts` or `custom_v1`). The dataset name is used as a dictionary key in [`config.json`](https://github.com/junyiye/textflow/blob/main/config.json) and as a folder name in the file system. Special characters or spaces may cause issues with the CLI argument parser or file path resolution in [`src/vqa.py`](https://github.com/junyiye/textflow/blob/main/src/vqa.py), [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), and [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py).