Difference Between `results_{timestamp}.json` and `eval_results_{timestamp}.jsonl` in twinkle-eval
The results_{timestamp}.json file contains a high-level summary of your entire evaluation run, while eval_results_{timestamp}.jsonl stores granular, per-question logs for detailed debugging and analysis.
When running evaluations with the ai-twinkle/eval framework (commonly invoked as twinkle-eval), the system generates two distinct output files that serve different analytical purposes. Understanding the difference between these files is essential for efficiently extracting aggregate metrics versus debugging individual model failures.
What Is results_{timestamp}.json?
The results_{timestamp}.json file serves as the high-level summary of your complete evaluation run. It aggregates performance metrics across all datasets and source files, providing a concise overview suitable for reporting and quick analysis.
Contents and Structure
This JSON file contains:
- Run metadata: Timestamp, configuration settings, and environment information
- Dataset-level aggregates: Average accuracy and standard deviation for each evaluated dataset
- File-level summaries: For each source file processed, it stores mean accuracy, standard deviation, and a reference path to the detailed log file
How It’s Generated
The file is produced by the JSON Exporter class located in twinkle_eval/results_exporters.py. The exporter enriches the raw final_results dictionary with environment data before serializing it:
enhanced_results = self._enhance_with_environment(results)
with open(output_path, "w", encoding="utf-8") as f:
json.dump(enhanced_results, f, indent=4, ensure_ascii=False)
(Source: twinkle_eval/results_exporters.py, lines 33-44)
What Is eval_results_{timestamp}.jsonl?
The eval_results_{timestamp}.jsonl file (note the .jsonl extension indicating JSON Lines format) contains the granular, per-question evaluation logs. Unlike the summary file, this captures every individual interaction between the model and the evaluation dataset.
Row-Level Question Data
Each line in the .jsonl file represents a single evaluated question and contains:
- Question metadata:
question_id,questiontext, andcorrect_answer - Model outputs: The LLM’s raw output and reasoning chain
- Evaluation metrics: Predicted answer, boolean correctness flag, and token usage statistics
Generation in the Evaluator
This file is generated inside the evaluator module at twinkle_eval/evaluators.py. After processing each source file, the evaluator writes detailed results in JSON Lines format:
results_path = os.path.join(results_dir, f"eval_results_{timestamp}.jsonl")
with open(results_path, "w", encoding="utf-8") as f:
for detail in detailed_results:
f.write(json.dumps(detail, ensure_ascii=False) + "\n")
(Source: twinkle_eval/evaluators.py, lines 37-41)
Key Differences Between the Two Files
Understanding when to use each file depends on their structural and functional distinctions:
-
Format:
results_{timestamp}.jsonis a single JSON object suitable for loading entirely into memory, whileeval_results_{timestamp}.jsonluses the JSON Lines format where each line is an independent JSON object, enabling streaming processing of large datasets. -
Granularity: The
.jsonfile provides aggregated statistics (means, standard deviations) at the dataset and file level. The.jsonlfile contains raw, unaggregated data for every single question evaluated. -
Use cases: Use the
.jsonfile for generating executive summaries, comparing model versions at a high level, or importing into spreadsheet tools. Use the.jsonlfile for debugging specific failures, analyzing per-question latency, or training error analysis models.
How the Files Work Together
The two output files are designed to complement each other through explicit cross-referencing. Within results_{timestamp}.json, the individual_runs field contains a list of paths pointing to the corresponding eval_results_{timestamp}.jsonl files:
"individual_runs": {
"results": ["results/eval_results_20240101_1200.jsonl"],
"accuracies": [0.87, 0.91]
}
This linkage enables the HTML exporter (and other downstream tools) to locate and embed granular details when generating interactive reports. The HTML exporter specifically reads the referenced .jsonl files to populate per-question visualizations while using the summary .json file for high-level dashboard metrics.
Practical Usage Examples
Loading the High-Level Summary
To quickly extract aggregate metrics from your evaluation run:
import json
with open("results/results_20240101_1200.json") as f:
summary = json.load(f)
print(f"Overall accuracy: {summary['dataset_averages']['accuracy']}")
print(f"Number of datasets evaluated: {len(summary['datasets'])}")
Streaming the Detailed Logs
For memory-efficient processing of individual question results (essential for large evaluations):
import json
details = []
with open("results/eval_results_20240101_1200.jsonl") as f:
for line in f:
details.append(json.loads(line))
# Filter for incorrect answers only
failures = [d for d in details if not d['is_correct']]
print(f"Total failures: {len(failures)}")
Correlating Both Files
To perform file-level analysis using both outputs:
import json
import os
# Load summary
with open("results/results_20240101_1200.json") as f:
summary = json.load(f)
# Access detailed logs for a specific dataset
for dataset_name, dataset_info in summary['datasets'].items():
jsonl_path = dataset_info['detailed_results_path']
if os.path.exists(jsonl_path):
with open(jsonl_path) as f:
questions = [json.loads(line) for line in f]
print(f"{dataset_name}: {len(questions)} questions evaluated")
Summary
-
results_{timestamp}.jsonprovides a high-level summary of the entire evaluation run, including aggregated accuracy metrics and configuration details, generated bytwinkle_eval/results_exporters.py. -
eval_results_{timestamp}.jsonlcontains granular, per-question logs in JSON Lines format, capturing every individual model interaction and correctness determination, generated bytwinkle_eval/evaluators.py. -
The summary file references the detailed JSONL files, enabling tools like the HTML exporter to combine high-level dashboards with drill-down question details.
-
Use the
.jsonfile for quick reporting and the.jsonlfile for deep debugging and failure analysis.
Frequently Asked Questions
Why does one file use .json and the other use .jsonl?
The .json extension indicates a single JSON document that must be parsed entirely into memory, which is appropriate for the compact summary data. The .jsonl (JSON Lines) format allows the detailed per-question logs to be streamed line-by-line without loading the entire file into RAM, which is essential when processing thousands of individual evaluation questions.
Can I run an evaluation and only generate one of these files?
While the framework generates both by default during a standard evaluation run, you can control output behavior through exporters. The results_{timestamp}.json is produced by the JSON Exporter (twinkle_eval/results_exporters.py), while the eval_results_{timestamp}.jsonl is written directly by the evaluator core. To suppress the detailed logs, you would need to modify the evaluator logic or post-process to delete the .jsonl files, though this is not recommended as it breaks HTML export functionality.
How do I correlate a specific question failure with the high-level summary?
First, locate the dataset and file containing the failure in the results_{timestamp}.json summary under the individual_runs or datasets section. Note the detailed_results_path field, which points to the specific eval_results_{timestamp}.jsonl file. Then stream that JSONL file and filter for entries where is_correct is false to identify the specific question_id and question text that failed.
Does the timestamp in both filenames always match?
Yes, both files use the same run timestamp generated at the start of the evaluation session. This synchronization ensures that the results_{timestamp}.json can correctly reference its corresponding eval_results_{timestamp}.jsonl files. If you run multiple evaluations, each will generate a new timestamped pair, preventing overwrites and maintaining clear audit trails.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →