TextFlow Evaluation Methodology: How to Assess Generated Answers and Interpret Results

TextFlow evaluates generated answers by treating assessment itself as a question-answering task, using a three-vote majority system with an LLM evaluator to determine per-sample accuracy and overall performance metrics.

The junyiye/textflow repository implements a rigorous evaluation methodology that transforms quality assessment into a structured reasoning task. Rather than relying on simple string matching, the system prompts a large language model to judge whether generated responses align with ground-truth answers, aggregating multiple evaluations to ensure robustness.

Understanding the TextFlow Evaluation Methodology

The evaluation process in src/evaluation.py treats every sample as a mini-question-answering problem. Instead of calculating token overlap or embedding similarity, the methodology asks an LLM to reason about semantic correctness.

Step 1: Building the Evaluation Prompt

The process begins by constructing a specialized prompt that provides the evaluator with full context. The load_evaluation_prompt function in src/prompts/prompts.py assembles a template containing:

  • The original question asked to the model
  • The generated response produced by TextFlow
  • The ground-truth answer from the dataset

This prompt frames the evaluation as a binary classification task, instructing the LLM to return either "Correct" or "Incorrect" based on semantic alignment rather than exact string matching.

Step 2: LLM Evaluation with Multiple Votes

To account for the stochastic nature of language models, the evaluation methodology calls the LLM evaluator three separate times for each sample. In src/evaluation.py (lines 58-63), the system iterates over three attempts, each time querying the default evaluator (gpt-4o) with the constructed prompt.

Each call returns a string judgment—"Correct" if the response aligns with the ground truth, "Incorrect" otherwise. This triple-vote approach mitigates the impact of occasional hallucinations or inconsistent reasoning by the evaluator model.

Step 3: Majority Vote Aggregation

Once three individual decisions are collected, the system aggregates them using a majority voting mechanism. The majority_vote function in src/utils.py (line 97) implements simple majority logic:

  • If two or three votes are "Correct", the final_decision is "Correct"
  • If two or three votes are "Incorrect", the final_decision is "Incorrect"

This consensus-based approach provides a more stable binary label than any single evaluation, reducing variance introduced by temperature sampling or prompt sensitivity in the evaluator LLM.

Step 4: Accuracy Calculation

After processing all samples, the evaluation script computes the final performance metric. In src/evaluation.py (lines 78-86), the system calculates:

accuracy = correct_count / total_count

Where correct_count represents the number of samples where final_decision equals "Correct". The script logs this percentage to the console, providing the primary quantitative measure of TextFlow's performance on the evaluated dataset.

How to Interpret TextFlow Evaluation Results

Understanding the output requires familiarity with the JSON structure produced by the evaluation pipeline. Each sample in the results file contains both individual vote records and the aggregated consensus.

Key Fields in the Output JSON

When you inspect an evaluated results file, look for these critical fields:

  • decision1, decision2, decision3: The raw string outputs from each of the three LLM evaluation calls. These may differ from one another due to sampling randomness.
  • final_decision: The majority-vote consensus derived from the three decisions above. This field contains either "Correct" or "Incorrect" and determines whether the sample counts toward the accuracy numerator.
  • accuracy: This value appears in the evaluation logs rather than individual JSON records, representing the aggregate percentage of correct samples across the entire dataset.

Understanding Accuracy Metrics

The accuracy percentage reported by TextFlow serves as the primary indicator of pipeline quality:

  • High accuracy (>80%): Indicates that the TextFlow pipeline (combining the textualizer and reasoner components) consistently produces answers that semantically align with ground truth on the tested dataset.
  • Moderate accuracy (60-80%): Suggests reasonable performance but indicates room for improvement in either the retrieval mechanism or the reasoning component.
  • Low accuracy (<60%): Signals significant mismatches between generated and ground-truth answers. In these cases, inspect individual final_decision fields to identify specific failure patterns.

When accuracy is lower than expected, trace back to specific samples by examining the decision1, decision2, and decision3 fields. If the three votes disagree frequently, the evaluation itself may be uncertain about the sample's quality, suggesting ambiguous ground truth or borderline correct responses.

Running the Evaluation and Inspecting Results

The evaluation pipeline is accessible via command-line execution and produces machine-readable JSON output that you can analyze programmatically.

Executing the Evaluation Script

To evaluate an experiment results file, run the evaluation module from the repository root:

python -m src.evaluation \
    --model_name gpt-4o \
    --data_path output/flowvqa/textflow/your_experiment.json

This command processes the specified JSON file, queries the LLM evaluator three times per sample, and logs the final accuracy to stdout. The script modifies the input file in-place (or writes to a specified output path) by adding the decision1, decision2, decision3, and final_decision fields to each record.

Inspecting Enriched JSON Data

After running the evaluation, you can programmatically inspect the results to understand per-sample performance:

import json

with open("output/flowvqa/textflow/your_experiment.json") as f:
    data = json.load(f)

# Print enriched records to inspect voting patterns

for i, (sample_id, record) in enumerate(data.items()):
    if i >= 3:
        break
    print(f"Sample {sample_id}:")
    print(f"  Question     : {record['question']}")
    print(f"  Response     : {record['response']}")
    print(f"  Ground-Truth : {record['answer']}")
    print(f"  Votes        : {record['decision1']}, {record['decision2']}, {record['decision3']}")
    print(f"  Final Verdict: {record['final_decision']}")

This inspection reveals whether disagreements between decision1, decision2, and decision3 correlate with specific question types or response patterns.

Computing Accuracy Manually

If you need to verify the reported accuracy or calculate it for a subset of samples:

correct = sum(
    1 for sample in data.values() 
    if sample.get("final_decision") == "Correct"
)
total = len(data)
accuracy = correct / total

print(f"Accuracy: {accuracy:.2%} ({correct}/{total})")

This manual calculation should match the accuracy logged by the evaluation script in src/evaluation.py.

Summary

TextFlow implements a robust evaluation methodology that treats quality assessment as a reasoning task rather than a mechanical comparison. Key takeaways include:

  • Triple-vote consensus: Each sample receives three independent judgments from an LLM evaluator (default gpt-4o) to mitigate stochastic variance.
  • Majority aggregation: The majority_vote function in src/utils.py converts three binary decisions into a single final_decision per sample.
  • Semantic evaluation: The load_evaluation_prompt function in src/prompts/prompts.py frames assessment as a correctness judgment rather than string matching, capturing semantic equivalence.
  • Accuracy calculation: Overall performance is computed in src/evaluation.py as the percentage of samples where final_decision equals "Correct".

Frequently Asked Questions

How does TextFlow handle ambiguous or borderline correct answers?

When the ground truth and generated response are semantically similar but not identical, the three LLM evaluators may disagree. TextFlow captures this uncertainty by storing all three votes (decision1, decision2, decision3) in the output JSON. If the votes split (e.g., two "Correct" and one "Incorrect"), the final_decision follows the majority, but the disagreement itself signals that the sample may be ambiguous. You can identify these cases by filtering for records where not all three decisions match.

Can I use a different LLM evaluator instead of GPT-4o?

Yes, the evaluation script accepts a --model_name parameter that allows you to specify alternative evaluators. While the default configuration uses gpt-4o for its strong reasoning capabilities, you can substitute other models supported by the ModelWrapper class in src/models.py. When changing evaluators, consider that the consistency of the three-vote system depends on the model's reasoning stability; weaker models may produce more variable decision1, decision2, and decision3 values, potentially requiring more than three votes for reliable aggregation.

Why does TextFlow use three votes instead of a single evaluation call?

The three-vote methodology addresses the inherent stochasticity of large language models. A single evaluation call might produce an anomalous judgment due to temperature sampling or prompt sensitivity. By querying the evaluator three times and applying the majority_vote function in src/utils.py, TextFlow achieves a more stable binary classification. This consensus approach reduces the variance in the final_decision field and ensures that the computed accuracy metric reflects genuine pipeline performance rather than evaluator randomness.

How do I debug low accuracy scores in my TextFlow experiments?

When the evaluation reports accuracy below your target threshold, inspect the enriched JSON output to identify failure patterns. Load the results file and filter for records where final_decision equals "Incorrect". Examine the decision1, decision2, and decision3 fields: unanimous "Incorrect" votes indicate clear errors in the TextFlow pipeline, while split votes suggest ambiguous ground truth. Cross-reference the question, response, and answer fields to determine whether errors stem from retrieval failures, reasoning mistakes, or ground-truth inconsistencies. This granular inspection allows you to trace accuracy issues to specific components of your experiment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →