How to Use Ground Truth Text Representations for Benchmarking in TextFlow
Pass --textualizer Gound-Truth to src/reasoner.py and provide a JSON file containing the perfect textual descriptions to benchmark your reasoning LLM against ideal inputs.
TextFlow (junyiye/textflow) supports benchmarking reasoning models against ground truth text representations instead of VLM-generated outputs. This workflow allows you to measure the upper-bound performance of your reasoning LLM when provided with perfect textual descriptions of flowcharts. By bypassing the visual-language model, you isolate and evaluate the pure reasoning capabilities of your language model using exact ground truth data.
Prerequisites: Preparing the Ground Truth File
Before running the benchmark, you must create a JSON file containing the ground truth textual representations. The reasoner expects this file at a specific path determined by your configuration.
Locate the output root directory defined in src/config.py under config["file_paths"]["output"]. Inside this directory, create the following file structure:
<output_root>/<dataset>/<input_type>/Gound-Truth.json
The JSON file must contain a dictionary mapping each sample key to its corresponding ground truth text representation. These representations can be in Mermaid, GraphViz, PlantUML, or any other supported format.
{
"0001": "graph TD;\n A --> B;\n B --> C;",
"0002": "graph TD;\n X --> Y;",
"0003": "flowchart LR\n Start --> Process --> End"
}
Ensure the sample keys match exactly with the dataset keys used in your evaluation split.
Running the Benchmark with Ground Truth
To execute the benchmark using ground truth representations, invoke src/reasoner.py with the special textualizer name Gound-Truth. Note that this spelling matches the implementation in the source code.
python -m src/reasoner \
--dataset flowvqa \
--reasoner Llama-3.1-8B \
--textualizer Gound-Truth \
--input_type mermaid
When you specify --textualizer Gound-Truth, the script directly loads the pre-existing textual representations from the JSON file rather than generating them from images. This flag signals the reasoner to skip the VLM inference step entirely and proceed directly to the reasoning phase.
How the Reasoner Processes Ground Truth Inputs
Inside src/reasoner.py, the workflow branches when the textualizer argument equals Gound-Truth. Instead of calling the visual-language model through src/textualizer.py, the reasoner loads the ground truth strings from the JSON file located at the constructed path.
The loaded text is then injected into the prompt template returned by load_reasoner_prompt (defined in src/prompts/prompts.py). The reasoning LLM—wrapped by the ModelWrapper class in src/models/model_loader.py—receives this perfect textual description and generates answers based solely on the ground truth input.
Results are written to the standard output file following the pattern *_reasoner_*.json within the output directory. This file contains the model's responses for each sample, ready for evaluation against gold standard answers.
Evaluating Benchmark Results
After generating predictions, compare the reasoner outputs against the ground truth answers located in the dataset's test.json file. TextFlow provides evaluation utilities in src/evaluation.py to compute accuracy and other benchmark metrics.
You can also implement custom evaluation scripts by loading the *_reasoner_*.json output and comparing it against the reference data. This step quantifies how well the reasoning LLM performs when visualization errors are eliminated from the pipeline.
Summary
- Prepare a
Gound-Truth.jsonfile mapping sample IDs to perfect textual representations and place it in<output_root>/<dataset>/<input_type>/. - Execute the benchmark by running
src/reasoner.pywith--textualizer Gound-Truthto bypass VLM generation. - Process involves the reasoner loading ground truth text via
load_reasoner_promptand querying the LLM throughModelWrapper. - Evaluate results using
src/evaluation.pyor custom scripts against the dataset's gold answers to measure pure reasoning performance.
Frequently Asked Questions
What file format does the ground truth need to be in?
The ground truth must be stored as a JSON file containing a flat dictionary where keys are sample identifiers and values are strings containing the textual representation (such as Mermaid or GraphViz syntax). Each value should be a properly escaped string representing the complete flowchart description.
Where should I place the Gound-Truth.json file?
Place the file at <output_root>/<dataset>/<input_type>/Gound-Truth.json, where output_root is defined in src/config.py under config["file_paths"]["output"]. The <dataset> and <input_type> directories must match the arguments passed to the reasoner script.
Does using ground truth skip the visual language model entirely?
Yes. When --textualizer Gound-Truth is specified, src/reasoner.py bypasses the VLM inference step completely. Instead of generating representations from images, it loads pre-existing text directly from the JSON file, allowing you to benchmark the reasoning LLM in isolation.
How do I evaluate the results after running the benchmark?
Use the utilities in src/evaluation.py to compare the generated answers in the *_reasoner_*.json output file against the gold standard answers in the dataset's test.json file. This comparison yields metrics that reflect the reasoning LLM's performance when provided with perfect visual descriptions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →