# TextFlow Evaluation Methodology: How to Assess Generated Answers and Interpret Results

> Discover the TextFlow evaluation methodology. Learn how TextFlow assesses generated answers using a unique LLM-powered majority vote system and interpret performance metrics effectively.

- Repository: [Junyi Ye/textflow](https://github.com/junyiye/textflow)
- Tags: evaluation-methodology
- Published: 2026-03-05

---

**TextFlow evaluates generated answers by treating assessment itself as a question-answering task, using a three-vote majority system with an LLM evaluator to determine per-sample accuracy and overall performance metrics.**

The `junyiye/textflow` repository implements a rigorous evaluation methodology that transforms quality assessment into a structured reasoning task. Rather than relying on simple string matching, the system prompts a large language model to judge whether generated responses align with ground-truth answers, aggregating multiple evaluations to ensure robustness.

## Understanding the TextFlow Evaluation Methodology

The evaluation process in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) treats every sample as a mini-question-answering problem. Instead of calculating token overlap or embedding similarity, the methodology asks an LLM to reason about semantic correctness.

### Step 1: Building the Evaluation Prompt

The process begins by constructing a specialized prompt that provides the evaluator with full context. The `load_evaluation_prompt` function in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) assembles a template containing:

- The original question asked to the model
- The generated response produced by TextFlow
- The ground-truth answer from the dataset

This prompt frames the evaluation as a binary classification task, instructing the LLM to return either `"Correct"` or `"Incorrect"` based on semantic alignment rather than exact string matching.

### Step 2: LLM Evaluation with Multiple Votes

To account for the stochastic nature of language models, the evaluation methodology calls the LLM evaluator three separate times for each sample. In [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) (lines 58-63), the system iterates over three attempts, each time querying the default evaluator (`gpt-4o`) with the constructed prompt.

Each call returns a string judgment—`"Correct"` if the response aligns with the ground truth, `"Incorrect"` otherwise. This triple-vote approach mitigates the impact of occasional hallucinations or inconsistent reasoning by the evaluator model.

### Step 3: Majority Vote Aggregation

Once three individual decisions are collected, the system aggregates them using a majority voting mechanism. The `majority_vote` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) (line 97) implements simple majority logic:

- If two or three votes are `"Correct"`, the `final_decision` is `"Correct"`
- If two or three votes are `"Incorrect"`, the `final_decision` is `"Incorrect"`

This consensus-based approach provides a more stable binary label than any single evaluation, reducing variance introduced by temperature sampling or prompt sensitivity in the evaluator LLM.

### Step 4: Accuracy Calculation

After processing all samples, the evaluation script computes the final performance metric. In [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) (lines 78-86), the system calculates:

```python
accuracy = correct_count / total_count

```

Where `correct_count` represents the number of samples where `final_decision` equals `"Correct"`. The script logs this percentage to the console, providing the primary quantitative measure of TextFlow's performance on the evaluated dataset.

## How to Interpret TextFlow Evaluation Results

Understanding the output requires familiarity with the JSON structure produced by the evaluation pipeline. Each sample in the results file contains both individual vote records and the aggregated consensus.

### Key Fields in the Output JSON

When you inspect an evaluated results file, look for these critical fields:

- **`decision1`**, **`decision2`**, **`decision3`**: The raw string outputs from each of the three LLM evaluation calls. These may differ from one another due to sampling randomness.
- **`final_decision`**: The majority-vote consensus derived from the three decisions above. This field contains either `"Correct"` or `"Incorrect"` and determines whether the sample counts toward the accuracy numerator.
- **`accuracy`**: This value appears in the evaluation logs rather than individual JSON records, representing the aggregate percentage of correct samples across the entire dataset.

### Understanding Accuracy Metrics

The accuracy percentage reported by TextFlow serves as the primary indicator of pipeline quality:

- **High accuracy (>80%)**: Indicates that the TextFlow pipeline (combining the textualizer and reasoner components) consistently produces answers that semantically align with ground truth on the tested dataset.
- **Moderate accuracy (60-80%)**: Suggests reasonable performance but indicates room for improvement in either the retrieval mechanism or the reasoning component.
- **Low accuracy (<60%)**: Signals significant mismatches between generated and ground-truth answers. In these cases, inspect individual `final_decision` fields to identify specific failure patterns.

When accuracy is lower than expected, trace back to specific samples by examining the `decision1`, `decision2`, and `decision3` fields. If the three votes disagree frequently, the evaluation itself may be uncertain about the sample's quality, suggesting ambiguous ground truth or borderline correct responses.

## Running the Evaluation and Inspecting Results

The evaluation pipeline is accessible via command-line execution and produces machine-readable JSON output that you can analyze programmatically.

### Executing the Evaluation Script

To evaluate an experiment results file, run the evaluation module from the repository root:

```bash
python -m src.evaluation \
    --model_name gpt-4o \
    --data_path output/flowvqa/textflow/your_experiment.json

```

This command processes the specified JSON file, queries the LLM evaluator three times per sample, and logs the final accuracy to stdout. The script modifies the input file in-place (or writes to a specified output path) by adding the `decision1`, `decision2`, `decision3`, and `final_decision` fields to each record.

### Inspecting Enriched JSON Data

After running the evaluation, you can programmatically inspect the results to understand per-sample performance:

```python
import json

with open("output/flowvqa/textflow/your_experiment.json") as f:
    data = json.load(f)

# Print enriched records to inspect voting patterns

for i, (sample_id, record) in enumerate(data.items()):
    if i >= 3:
        break
    print(f"Sample {sample_id}:")
    print(f"  Question     : {record['question']}")
    print(f"  Response     : {record['response']}")
    print(f"  Ground-Truth : {record['answer']}")
    print(f"  Votes        : {record['decision1']}, {record['decision2']}, {record['decision3']}")
    print(f"  Final Verdict: {record['final_decision']}")

```

This inspection reveals whether disagreements between `decision1`, `decision2`, and `decision3` correlate with specific question types or response patterns.

### Computing Accuracy Manually

If you need to verify the reported accuracy or calculate it for a subset of samples:

```python
correct = sum(
    1 for sample in data.values() 
    if sample.get("final_decision") == "Correct"
)
total = len(data)
accuracy = correct / total

print(f"Accuracy: {accuracy:.2%} ({correct}/{total})")

```

This manual calculation should match the accuracy logged by the evaluation script in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py).

## Summary

TextFlow implements a robust **evaluation methodology** that treats quality assessment as a reasoning task rather than a mechanical comparison. Key takeaways include:

- **Triple-vote consensus**: Each sample receives three independent judgments from an LLM evaluator (default `gpt-4o`) to mitigate stochastic variance.
- **Majority aggregation**: The `majority_vote` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py) converts three binary decisions into a single `final_decision` per sample.
- **Semantic evaluation**: The `load_evaluation_prompt` function in [`src/prompts/prompts.py`](https://github.com/junyiye/textflow/blob/main/src/prompts/prompts.py) frames assessment as a correctness judgment rather than string matching, capturing semantic equivalence.
- **Accuracy calculation**: Overall performance is computed in [`src/evaluation.py`](https://github.com/junyiye/textflow/blob/main/src/evaluation.py) as the percentage of samples where `final_decision` equals `"Correct"`.

## Frequently Asked Questions

### How does TextFlow handle ambiguous or borderline correct answers?

When the ground truth and generated response are semantically similar but not identical, the three LLM evaluators may disagree. TextFlow captures this uncertainty by storing all three votes (`decision1`, `decision2`, `decision3`) in the output JSON. If the votes split (e.g., two "Correct" and one "Incorrect"), the `final_decision` follows the majority, but the disagreement itself signals that the sample may be ambiguous. You can identify these cases by filtering for records where not all three decisions match.

### Can I use a different LLM evaluator instead of GPT-4o?

Yes, the evaluation script accepts a `--model_name` parameter that allows you to specify alternative evaluators. While the default configuration uses `gpt-4o` for its strong reasoning capabilities, you can substitute other models supported by the `ModelWrapper` class in [`src/models.py`](https://github.com/junyiye/textflow/blob/main/src/models.py). When changing evaluators, consider that the consistency of the three-vote system depends on the model's reasoning stability; weaker models may produce more variable `decision1`, `decision2`, and `decision3` values, potentially requiring more than three votes for reliable aggregation.

### Why does TextFlow use three votes instead of a single evaluation call?

The three-vote methodology addresses the inherent stochasticity of large language models. A single evaluation call might produce an anomalous judgment due to temperature sampling or prompt sensitivity. By querying the evaluator three times and applying the `majority_vote` function in [`src/utils.py`](https://github.com/junyiye/textflow/blob/main/src/utils.py), TextFlow achieves a more stable binary classification. This consensus approach reduces the variance in the `final_decision` field and ensures that the computed `accuracy` metric reflects genuine pipeline performance rather than evaluator randomness.

### How do I debug low accuracy scores in my TextFlow experiments?

When the evaluation reports accuracy below your target threshold, inspect the enriched JSON output to identify failure patterns. Load the results file and filter for records where `final_decision` equals `"Incorrect"`. Examine the `decision1`, `decision2`, and `decision3` fields: unanimous "Incorrect" votes indicate clear errors in the TextFlow pipeline, while split votes suggest ambiguous ground truth. Cross-reference the `question`, `response`, and `answer` fields to determine whether errors stem from retrieval failures, reasoning mistakes, or ground-truth inconsistencies. This granular inspection allows you to trace accuracy issues to specific components of your experiment.