How Twinkle Eval Implements Multi-Run Stability Analysis for LLM Evaluation
Twinkle Eval implements multi-run stability analysis through the repeat_runs configuration option, which executes each evaluation file multiple times and aggregates results using NumPy to calculate mean accuracy and standard deviation.
The ai-twinkle/eval repository provides a robust framework for assessing large language model performance. Its multi-run stability analysis feature allows developers to measure how consistently a model performs across repeated evaluations, addressing the inherent variability in LLM outputs.
Understanding the Multi-Run Stability Analysis Workflow
The stability analysis operates through a three-stage pipeline: configuration validation, repeated execution, and statistical aggregation. This design ensures that users can quantify performance variance while maintaining clean separation between execution logic and result analysis.
Step 1: Configuring Repeat Runs in Twinkle Eval
Setting the repeat_runs Parameter
Users enable multi-run analysis by setting the evaluation.repeat_runs value in their YAML configuration file. The default value is 1, meaning single-run execution occurs unless explicitly overridden.
evaluation:
repeat_runs: 5 # Execute each file 5 times
shuffle_options: false
datasets_prompt_map: {}
dataset_paths:
- ./datasets/mmlu
Validation in validators.py
Before execution begins, the configuration undergoes validation in twinkle_eval/validators.py. Lines 51-55 ensure that repeat_runs is a positive integer, preventing invalid configurations from reaching the execution stage.
# From twinkle_eval/validators.py (lines 51-55)
if not isinstance(config.repeat_runs, int) or config.repeat_runs < 1:
raise ValidationError(
f"repeat_runs must be a positive integer, got {config.repeat_runs}"
)
Step 2: Executing Multiple Evaluation Runs
The _evaluate_dataset Method
The core execution logic resides in TwinkleEvalRunner._evaluate_dataset within twinkle_eval/main.py. Lines 200-207 implement the iteration logic that re-executes each evaluation file according to the configured repeat_runs value.
For every iteration, the runner invokes Evaluator.evaluate_file, captures the accuracy score, and records the path to the detailed result file. This collection phase creates the raw data necessary for statistical analysis.
# Conceptual implementation from twinkle_eval/main.py (lines 200-207)
accuracies = []
result_files = []
for run_idx in range(config.repeat_runs):
accuracy, result_path = evaluator.evaluate_file(
file_path,
run_index=run_idx
)
accuracies.append(accuracy)
result_files.append(result_path)
Step 3: Aggregating Stability Metrics with NumPy
Computing Mean and Standard Deviation
After collecting all per-run accuracies, the system calculates stability metrics using NumPy operations in twinkle_eval/main.py (lines 224-238). The implementation computes both the mean accuracy (np.mean) and standard deviation (np.std) across all runs.
These metrics quantify central tendency and variance, allowing users to distinguish between consistently high-performing models and those with erratic behavior.
# From twinkle_eval/main.py (lines 224-238)
import numpy as np
accuracy_mean = np.mean(accuracies)
accuracy_std = np.std(accuracies)
file_result = {
"file": file_path,
"accuracy_mean": float(accuracy_mean),
"accuracy_std": float(accuracy_std),
"individual_runs": {
"accuracies": accuracies,
"results": result_files
}
}
Output Structure
The aggregated results follow a structured JSON format that preserves both summary statistics and granular per-run data. Each evaluated file receives an entry containing the mean accuracy, standard deviation, and arrays of individual run results.
{
"file": "path/to/example.py",
"accuracy_mean": 0.92,
"accuracy_std": 0.03,
"individual_runs": {
"accuracies": [0.90, 0.94, 0.92],
"results": ["results/..._run0.json", "results/..._run1.json", "results/..._run2.json"]
}
}
Practical Implementation Examples
YAML Configuration
Enable stability analysis by specifying repeat_runs in your evaluation configuration:
evaluation:
repeat_runs: 5 # run each file 5 times
shuffle_options: false
datasets_prompt_map: {}
dataset_paths:
- ./datasets/mmlu
Command Line Execution
Run the evaluation from the CLI to automatically trigger multi-run analysis:
twinkle-eval --config my_config.yaml --export json html
The runner executes five evaluations per file, computes mean and standard deviation, and exports the enriched results.
Programmatic Access
Access stability metrics directly through the Python API:
from twinkle_eval.main import TwinkleEvalRunner
runner = TwinkleEvalRunner("my_config.yaml")
runner.initialize()
final_json = runner.run_evaluation(export_formats=["json"])
# Load the generated JSON
import json, pathlib
with open(pathlib.Path(final_json).with_suffix('.json'), "r", encoding="utf-8") as f:
results = json.load(f)
# Print stability metrics for a specific file
file_stats = results["dataset_results"]["./datasets/mmlu"]["results"][0]
print("Mean:", file_stats["accuracy_mean"])
print("Std :", file_stats["accuracy_std"])
Key Source Files for Multi-Run Stability Analysis
twinkle_eval/config.py— Supplies the defaultrepeat_runsvalue (set to1) and defines the configuration schema.twinkle_eval/validators.py— Validates thatrepeat_runsis a positive integer before execution begins (lines 51-55).twinkle_eval/main.py— Implements the core execution loop inTwinkleEvalRunner._evaluate_dataset(lines 200-207) and the statistical aggregation using NumPy (lines 224-238).twinkle_eval/results_exporters.py— Persists the aggregated stability data into JSON, HTML, and Excel formats.
Summary
- Twinkle Eval implements multi-run stability analysis through the
repeat_runsconfiguration option, allowing users to execute each evaluation file multiple times. - The
TwinkleEvalRunner._evaluate_datasetmethod intwinkle_eval/main.pyorchestrates the repeated execution and collects per-run accuracy scores. - Statistical aggregation uses NumPy to calculate
accuracy_meanandaccuracy_std, providing quantitative measures of performance stability. - Validation occurs in
twinkle_eval/validators.pyto ensurerepeat_runsis a positive integer before execution. - Results include both summary statistics and granular per-run data, enabling detailed analysis of model consistency.
Frequently Asked Questions
What is the default value for repeat_runs in Twinkle Eval?
The default value for repeat_runs is 1, defined in twinkle_eval/config.py. This means single-run execution occurs unless the user explicitly configures a higher value in the YAML configuration file.
How does Twinkle Eval calculate stability metrics?
Twinkle Eval calculates stability metrics using NumPy operations in twinkle_eval/main.py (lines 224-238). After collecting accuracy scores from all runs, it computes the mean using np.mean() and the standard deviation using np.std(). These values are stored as accuracy_mean and accuracy_std in the results.
Can I access individual run results after aggregation?
Yes, the aggregated results preserve individual run data. The output JSON includes an individual_runs object containing two arrays: accuracies (the raw accuracy scores) and results (file paths to detailed per-run result files). This allows you to drill down into specific runs while still viewing summary statistics.
Where is the stability analysis logic implemented?
The core stability analysis logic resides in twinkle_eval/main.py within the TwinkleEvalRunner class. Specifically, lines 200-207 handle the repeated execution loop, while lines 224-238 perform the statistical aggregation. Configuration validation occurs in twinkle_eval/validators.py (lines 51-55), and result persistence is handled by twinkle_eval/results_exporters.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →