# How to Use Ground Truth Text Representations for Benchmarking in TextFlow

> Benchmark your LLM using ground truth text representations in TextFlow. Pass --textualizer Ground-Truth and provide ideal text descriptions for accurate evaluation.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Pass `--textualizer Gound-Truth` to [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) and provide a JSON file containing the perfect textual descriptions to benchmark your reasoning LLM against ideal inputs.**

TextFlow (junyiye/textflow) supports benchmarking reasoning models against **ground truth text representations** instead of VLM-generated outputs. This workflow allows you to measure the upper-bound performance of your reasoning LLM when provided with perfect textual descriptions of flowcharts. By bypassing the visual-language model, you isolate and evaluate the pure reasoning capabilities of your language model using exact ground truth data.

## Prerequisites: Preparing the Ground Truth File

Before running the benchmark, you must create a JSON file containing the ground truth textual representations. The reasoner expects this file at a specific path determined by your configuration.

Locate the output root directory defined in [`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py) under `config["file_paths"]["output"]`. Inside this directory, create the following file structure:

```

<output_root>/<dataset>/<input_type>/Gound-Truth.json

```

The JSON file must contain a dictionary mapping each sample key to its corresponding ground truth text representation. These representations can be in Mermaid, GraphViz, PlantUML, or any other supported format.

```json
{
    "0001": "graph TD;\n  A --> B;\n  B --> C;",
    "0002": "graph TD;\n  X --> Y;",
    "0003": "flowchart LR\n    Start --> Process --> End"
}

```

Ensure the sample keys match exactly with the dataset keys used in your evaluation split.

## Running the Benchmark with Ground Truth

To execute the benchmark using ground truth representations, invoke [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) with the special textualizer name `Gound-Truth`. Note that this spelling matches the implementation in the source code.

```bash
python -m src/reasoner \
    --dataset flowvqa \
    --reasoner Llama-3.1-8B \
    --textualizer Gound-Truth \
    --input_type mermaid

```

When you specify `--textualizer Gound-Truth`, the script directly loads the pre-existing textual representations from the JSON file rather than generating them from images. This flag signals the reasoner to skip the VLM inference step entirely and proceed directly to the reasoning phase.

## How the Reasoner Processes Ground Truth Inputs

Inside [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py), the workflow branches when the textualizer argument equals `Gound-Truth`. Instead of calling the visual-language model through [`src/textualizer.py`](https://github.com/junyiye/textflow/blob/main/src/textualizer.py), the reasoner loads the ground truth strings from the JSON file located at the constructed path.

The loaded text is then injected into the prompt template returned by `load_reasoner_prompt` (defined in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py)). The reasoning LLM—wrapped by the `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/textflow/blob/main/src/models/model_loader.py)—receives this perfect textual description and generates answers based solely on the ground truth input.

Results are written to the standard output file following the pattern `*_reasoner_*.json` within the output directory. This file contains the model's responses for each sample, ready for evaluation against gold standard answers.

## Evaluating Benchmark Results

After generating predictions, compare the reasoner outputs against the ground truth answers located in the dataset's [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) file. TextFlow provides evaluation utilities in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) to compute accuracy and other benchmark metrics.

You can also implement custom evaluation scripts by loading the `*_reasoner_*.json` output and comparing it against the reference data. This step quantifies how well the reasoning LLM performs when visualization errors are eliminated from the pipeline.

## Summary

- **Prepare** a [`Gound-Truth.json`](https://github.com/junyiye/textflow/blob/main/Gound-Truth.json) file mapping sample IDs to perfect textual representations and place it in `<output_root>/<dataset>/<input_type>/`.
- **Execute** the benchmark by running [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) with `--textualizer Gound-Truth` to bypass VLM generation.
- **Process** involves the reasoner loading ground truth text via `load_reasoner_prompt` and querying the LLM through `ModelWrapper`.
- **Evaluate** results using [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) or custom scripts against the dataset's gold answers to measure pure reasoning performance.

## Frequently Asked Questions

### What file format does the ground truth need to be in?

The ground truth must be stored as a JSON file containing a flat dictionary where keys are sample identifiers and values are strings containing the textual representation (such as Mermaid or GraphViz syntax). Each value should be a properly escaped string representing the complete flowchart description.

### Where should I place the Gound-Truth.json file?

Place the file at `<output_root>/<dataset>/<input_type>/Gound-Truth.json`, where `output_root` is defined in [`src/config.py`](https://github.com/junyiye/textflow/blob/main/src/config.py) under `config["file_paths"]["output"]`. The `<dataset>` and `<input_type>` directories must match the arguments passed to the reasoner script.

### Does using ground truth skip the visual language model entirely?

Yes. When `--textualizer Gound-Truth` is specified, [`src/reasoner.py`](https://github.com/junyiye/textflow/blob/main/src/reasoner.py) bypasses the VLM inference step completely. Instead of generating representations from images, it loads pre-existing text directly from the JSON file, allowing you to benchmark the reasoning LLM in isolation.

### How do I evaluate the results after running the benchmark?

Use the utilities in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) to compare the generated answers in the `*_reasoner_*.json` output file against the gold standard answers in the dataset's [`test.json`](https://github.com/junyiye/textflow/blob/main/test.json) file. This comparison yields metrics that reflect the reasoning LLM's performance when provided with perfect visual descriptions.