TextFlow Evaluation Methodology: How to Assess Generated Answers and Interpret Results
TextFlow evaluates generated answers by treating assessment itself as a question-answering task, using a three-vote majority system with an LLM evaluator to determine per-sample accuracy and overall performance metrics.
The junyiye/textflow repository implements a rigorous evaluation methodology that transforms quality assessment into a structured reasoning task. Rather than relying on simple string matching, the system prompts a large language model to judge whether generated responses align with ground-truth answers, aggregating multiple evaluations to ensure robustness.
Understanding the TextFlow Evaluation Methodology
The evaluation process in src/evaluation.py treats every sample as a mini-question-answering problem. Instead of calculating token overlap or embedding similarity, the methodology asks an LLM to reason about semantic correctness.
Step 1: Building the Evaluation Prompt
The process begins by constructing a specialized prompt that provides the evaluator with full context. The load_evaluation_prompt function in src/prompts/prompts.py assembles a template containing:
- The original question asked to the model
- The generated response produced by TextFlow
- The ground-truth answer from the dataset
This prompt frames the evaluation as a binary classification task, instructing the LLM to return either "Correct" or "Incorrect" based on semantic alignment rather than exact string matching.
Step 2: LLM Evaluation with Multiple Votes
To account for the stochastic nature of language models, the evaluation methodology calls the LLM evaluator three separate times for each sample. In src/evaluation.py (lines 58-63), the system iterates over three attempts, each time querying the default evaluator (gpt-4o) with the constructed prompt.
Each call returns a string judgment—"Correct" if the response aligns with the ground truth, "Incorrect" otherwise. This triple-vote approach mitigates the impact of occasional hallucinations or inconsistent reasoning by the evaluator model.
Step 3: Majority Vote Aggregation
Once three individual decisions are collected, the system aggregates them using a majority voting mechanism. The majority_vote function in src/utils.py (line 97) implements simple majority logic:
- If two or three votes are
"Correct", thefinal_decisionis"Correct" - If two or three votes are
"Incorrect", thefinal_decisionis"Incorrect"
This consensus-based approach provides a more stable binary label than any single evaluation, reducing variance introduced by temperature sampling or prompt sensitivity in the evaluator LLM.
Step 4: Accuracy Calculation
After processing all samples, the evaluation script computes the final performance metric. In src/evaluation.py (lines 78-86), the system calculates:
accuracy = correct_count / total_count
Where correct_count represents the number of samples where final_decision equals "Correct". The script logs this percentage to the console, providing the primary quantitative measure of TextFlow's performance on the evaluated dataset.
How to Interpret TextFlow Evaluation Results
Understanding the output requires familiarity with the JSON structure produced by the evaluation pipeline. Each sample in the results file contains both individual vote records and the aggregated consensus.
Key Fields in the Output JSON
When you inspect an evaluated results file, look for these critical fields:
decision1,decision2,decision3: The raw string outputs from each of the three LLM evaluation calls. These may differ from one another due to sampling randomness.final_decision: The majority-vote consensus derived from the three decisions above. This field contains either"Correct"or"Incorrect"and determines whether the sample counts toward the accuracy numerator.accuracy: This value appears in the evaluation logs rather than individual JSON records, representing the aggregate percentage of correct samples across the entire dataset.
Understanding Accuracy Metrics
The accuracy percentage reported by TextFlow serves as the primary indicator of pipeline quality:
- High accuracy (>80%): Indicates that the TextFlow pipeline (combining the textualizer and reasoner components) consistently produces answers that semantically align with ground truth on the tested dataset.
- Moderate accuracy (60-80%): Suggests reasonable performance but indicates room for improvement in either the retrieval mechanism or the reasoning component.
- Low accuracy (<60%): Signals significant mismatches between generated and ground-truth answers. In these cases, inspect individual
final_decisionfields to identify specific failure patterns.
When accuracy is lower than expected, trace back to specific samples by examining the decision1, decision2, and decision3 fields. If the three votes disagree frequently, the evaluation itself may be uncertain about the sample's quality, suggesting ambiguous ground truth or borderline correct responses.
Running the Evaluation and Inspecting Results
The evaluation pipeline is accessible via command-line execution and produces machine-readable JSON output that you can analyze programmatically.
Executing the Evaluation Script
To evaluate an experiment results file, run the evaluation module from the repository root:
python -m src.evaluation \
--model_name gpt-4o \
--data_path output/flowvqa/textflow/your_experiment.json
This command processes the specified JSON file, queries the LLM evaluator three times per sample, and logs the final accuracy to stdout. The script modifies the input file in-place (or writes to a specified output path) by adding the decision1, decision2, decision3, and final_decision fields to each record.
Inspecting Enriched JSON Data
After running the evaluation, you can programmatically inspect the results to understand per-sample performance:
import json
with open("output/flowvqa/textflow/your_experiment.json") as f:
data = json.load(f)
# Print enriched records to inspect voting patterns
for i, (sample_id, record) in enumerate(data.items()):
if i >= 3:
break
print(f"Sample {sample_id}:")
print(f" Question : {record['question']}")
print(f" Response : {record['response']}")
print(f" Ground-Truth : {record['answer']}")
print(f" Votes : {record['decision1']}, {record['decision2']}, {record['decision3']}")
print(f" Final Verdict: {record['final_decision']}")
This inspection reveals whether disagreements between decision1, decision2, and decision3 correlate with specific question types or response patterns.
Computing Accuracy Manually
If you need to verify the reported accuracy or calculate it for a subset of samples:
correct = sum(
1 for sample in data.values()
if sample.get("final_decision") == "Correct"
)
total = len(data)
accuracy = correct / total
print(f"Accuracy: {accuracy:.2%} ({correct}/{total})")
This manual calculation should match the accuracy logged by the evaluation script in src/evaluation.py.
Summary
TextFlow implements a robust evaluation methodology that treats quality assessment as a reasoning task rather than a mechanical comparison. Key takeaways include:
- Triple-vote consensus: Each sample receives three independent judgments from an LLM evaluator (default
gpt-4o) to mitigate stochastic variance. - Majority aggregation: The
majority_votefunction insrc/utils.pyconverts three binary decisions into a singlefinal_decisionper sample. - Semantic evaluation: The
load_evaluation_promptfunction insrc/prompts/prompts.pyframes assessment as a correctness judgment rather than string matching, capturing semantic equivalence. - Accuracy calculation: Overall performance is computed in
src/evaluation.pyas the percentage of samples wherefinal_decisionequals"Correct".
Frequently Asked Questions
How does TextFlow handle ambiguous or borderline correct answers?
When the ground truth and generated response are semantically similar but not identical, the three LLM evaluators may disagree. TextFlow captures this uncertainty by storing all three votes (decision1, decision2, decision3) in the output JSON. If the votes split (e.g., two "Correct" and one "Incorrect"), the final_decision follows the majority, but the disagreement itself signals that the sample may be ambiguous. You can identify these cases by filtering for records where not all three decisions match.
Can I use a different LLM evaluator instead of GPT-4o?
Yes, the evaluation script accepts a --model_name parameter that allows you to specify alternative evaluators. While the default configuration uses gpt-4o for its strong reasoning capabilities, you can substitute other models supported by the ModelWrapper class in src/models.py. When changing evaluators, consider that the consistency of the three-vote system depends on the model's reasoning stability; weaker models may produce more variable decision1, decision2, and decision3 values, potentially requiring more than three votes for reliable aggregation.
Why does TextFlow use three votes instead of a single evaluation call?
The three-vote methodology addresses the inherent stochasticity of large language models. A single evaluation call might produce an anomalous judgment due to temperature sampling or prompt sensitivity. By querying the evaluator three times and applying the majority_vote function in src/utils.py, TextFlow achieves a more stable binary classification. This consensus approach reduces the variance in the final_decision field and ensures that the computed accuracy metric reflects genuine pipeline performance rather than evaluator randomness.
How do I debug low accuracy scores in my TextFlow experiments?
When the evaluation reports accuracy below your target threshold, inspect the enriched JSON output to identify failure patterns. Load the results file and filter for records where final_decision equals "Incorrect". Examine the decision1, decision2, and decision3 fields: unanimous "Incorrect" votes indicate clear errors in the TextFlow pipeline, while split votes suggest ambiguous ground truth. Cross-reference the question, response, and answer fields to determine whether errors stem from retrieval failures, reasoning mistakes, or ground-truth inconsistencies. This granular inspection allows you to trace accuracy issues to specific components of your experiment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →