Where to Find Caveman's Benchmark Evaluation Results

Caveman stores its complete benchmark evaluation results in evals/snapshots/results.json, a JSON file containing timing, accuracy, and resource-usage metrics for each model configuration.

The JuliusBrussee/caveman repository includes a comprehensive benchmark suite for evaluating model performance. Understanding where to find Caveman's benchmark evaluation results is essential for analyzing model configurations and comparing performance metrics across different runs.

Primary Benchmark Results Location

The definitive source for benchmark outcomes is evals/snapshots/results.json. This file aggregates all measured scores from the test suite, including timing data, accuracy measurements, and resource-usage statistics for each evaluated model configuration.

You can view this file directly on GitHub at the following URL: https://github.com/JuliusBrussee/caveman/blob/main/evals/snapshots/results.json.

Exploring the Benchmarks Directory Structure

Current Results Storage

According to the source code, the evals/snapshots/results.json file serves as the primary artifact for benchmark data. This JSON structure contains a complete dump of benchmark runs, with each entry typically including the model identifier, specific metric names, and their corresponding values.

Future Results Artifacts

The repository also contains benchmarks/results/, a placeholder directory currently holding only a .gitkeep file. This directory is intended for future result artifacts such as CSV exports or visualization plots that may be generated from the raw JSON data.

How to Access and Parse the Results

Viewing Results on GitHub

For quick inspection without cloning, navigate directly to the results file in the browser. The JSON structure is human-readable and contains the full dataset of benchmark measurements.

Loading Results Programmatically

To analyze the benchmark data locally, clone the repository and use Python to parse the JSON file:

import json
from pathlib import Path

# Path to the results file (relative to the repository root)

results_path = Path(__file__).parent.parent / "evals" / "snapshots" / "results.json"

with results_path.open("r", encoding="utf-8") as f:
    results = json.load(f)

# Pretty-print a summary

for entry in results:
    model = entry.get("model")
    metric = entry.get("metric")
    value = entry.get("value")
    print(f"{model} – {metric}: {value}")

Inspecting the Results Directory

To verify the contents of the placeholder results directory, use the following command:


# List the contents of the benchmark results directory

ls -l benchmarks/results

Key Files in the Benchmark Pipeline

The benchmark evaluation system in JuliusBrussee/caveman consists of several interconnected components:

These files constitute the complete pipeline: run.py executes tests using prompts from prompts.json, dependencies are managed via requirements.txt, and final scores are captured in evals/snapshots/results.json.

Summary

  • Caveman's benchmark evaluation results are stored in evals/snapshots/results.json as a structured JSON file
  • The file contains timing, accuracy, and resource-usage metrics for each model configuration
  • benchmarks/results/ serves as a placeholder for future CSV or visualization artifacts
  • Use benchmarks/run.py to execute new benchmarks and regenerate results
  • Results can be accessed directly via GitHub or loaded programmatically using standard JSON parsing libraries

Frequently Asked Questions

Where is the main Caveman benchmark results file located?

The primary benchmark results file is located at evals/snapshots/results.json in the repository root. This JSON file contains a complete dump of all benchmark runs including timing, accuracy, and resource-usage metrics for each model configuration tested.

What metrics are included in Caveman's benchmark evaluation results?

According to the source code analysis, the results.json file includes timing measurements, accuracy scores, and resource-usage statistics. Each entry in the JSON array typically contains fields for the model identifier, specific metric names, and their corresponding numerical values.

Can I run the benchmarks locally to generate new results?

Yes. Execute benchmarks/run.py to run the benchmark suite locally. Ensure you install dependencies listed in benchmarks/requirements.txt first. The script will generate new measurements that can be compared against the existing evals/snapshots/results.json baseline.

Is there a visualization of the benchmark results?

Currently, the benchmarks/results/ directory contains only a .gitkeep file and serves as a placeholder for future artifacts. While the raw JSON data in evals/snapshots/results.json is machine-readable, you must generate your own visualizations or CSV exports from this data until official plotting tools are added to the repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →