How Harvey-Labs Comparison Dashboards Generate Aggregate Metrics Across Multiple Runs

Harvey-Labs aggregates evaluation runs by scanning the results/ directory for score files, deduplicating by timestamp, grouping runs by model label, and computing pooled and macro pass rates to power its comparison dashboards.

The harvey-labs repository provides a comprehensive evaluation framework for legal AI models, with comparison dashboards that synthesize results from dozens or hundreds of individual test runs. Understanding how harvey-labs comparison dashboards generate aggregate metrics across multiple runs reveals the pipeline that transforms raw JSON scores into leaderboard-ready statistics.

Scanning and Normalizing Individual Runs

The aggregation process begins in evaluation/compare.py, where the collect_runs() function (lines 52-68) walks the results/ directory tree to discover evaluation artifacts. It identifies two distinct file types: scores.json for single-judge evaluations and scores_dual.json for dual-judge setups.

Each discovered file passes through _comparison_scores(), which standardizes the raw JSON into a common schema. For dual-judge runs, this helper delegates to _normalize_dual_scores() imported from evaluation/report.py (lines 11-30). This normalization ensures that fields like model id, effort, token usage, cost, and pass/fail counts are uniformly accessible regardless of the evaluation mode. The extracted data accumulates in a temporary list named raw_runs before moving to the next stage.

Deduplicating Repeated Evaluations

When a model undergoes multiple evaluations against the same task, the system retains only the newest run to prevent skewed aggregates. This deduplication logic (lines 16-23 in evaluation/compare.py) compares timestamp-named directories and filters the raw_runs list to keep the latest entry per model-task combination.

Aggregating Metrics Across Tasks

The core mathematical aggregation happens in _aggregate_across_tasks() (lines 26-41), which collapses the flat list of runs into per-model summary statistics.

Grouping by Model Label

Runs cluster by their human-readable pretty_label, which encodes the model identifier and effort level. This grouping ensures that variants of the same model are analyzed separately while all tasks for a specific configuration roll up together.

Computing Pass Rates and Costs

For each model group, the function tallies cumulative values:

  • total_passed and total_criteria for raw accuracy counts
  • total_tokens, total_wall_clock, and total_cost for resource consumption
  • criterion_pass_fraction_sum for per-task pass ratios
  • all_pass_points for dual-judge agreement scores (which may be fractional)

From these sums, the system derives three critical metrics. The pooled criterion-pass rate (lines 84-89) divides total passes by total criteria across all tasks. The macro criterion-pass rate (lines 90-94) averages the per-task pass fractions, giving equal weight to each task regardless of its criteria count.

Deriving the All-Pass Score

The final leaderboard metric, the all-pass rate, originates from the summed all_pass_points (lines 96-101). For legacy single-judge runs, this value casts to an integer, while dual-judge runs maintain fractional agreement weights. The rate equals all_pass_count / n_tasks (lines 102-104). The function also tracks all_pass_both_agree_* counters to record judge concordance.

After calculation, the resulting dictionaries sort by the overall score (the all-pass rate) before returning to the caller (lines 138-140).

Rendering the Dashboard Visualizations

With aggregated data prepared, the compare_area() and compare_all() commands orchestrate visualization. These functions invoke the collection and aggregation pipeline, then hand the processed dictionaries to charting helpers in evaluation/charts.py, including leaderboard_table(), grouped_bars(), and rubric_vs_allpass_bars().

Matplotlib figures generate as either inline displays or PNG files. The _write_html() function (lines 31-38) embeds these assets into static HTML reports, producing the final comparison dashboards stored in results/comparisons/.

Practical Usage Examples

Generate a complete dashboard for a specific practice area using the high-level API:

from evaluation.compare import compare_area

# Produce HTML report for funds asset management evaluations

compare_area(area="funds-asset-management", save_images=False)

For custom reporting or programmatic analysis, access the aggregation pipeline directly:

from evaluation.compare import collect_runs, _aggregate_across_tasks

# Gather all runs from results/

runs = collect_runs()

# Define task scope

task_list = sorted({run["task"] for run in runs})

# Generate aggregates

aggregated = _aggregate_across_tasks(runs=runs, task_list=task_list)

# Access metrics for the top-performing model

top_model = aggregated[0]
print(f"All-pass rate: {top_model['all_pass_rate']}")
print(f"Pooled criterion rate: {top_model['criterion_pass_rate_pooled']}")
print(f"Total cost: ${top_model['total_cost']:.2f}")

Summary

  • Collection: collect_runs() in evaluation/compare.py scans for scores.json and scores_dual.json files, normalizing dual-judge outputs via evaluation/report.py.
  • Deduplication: Only the newest run per model-task pair survives, determined by timestamp directory naming.
  • Aggregation: _aggregate_across_tasks() groups by pretty_label and computes pooled rates (total passes / total criteria), macro rates (average of per-task fractions), and the all-pass rate (primary leaderboard metric).
  • Visualization: evaluation/charts.py consumes aggregated dictionaries to render tables and charts, with HTML generation handled by _write_html().
  • Entry Points: Use compare_area() for specific domains or compare_all() for comprehensive benchmarks.

Frequently Asked Questions

How does harvey-labs handle duplicate runs of the same model?

The system deduplicates by retaining only the newest run for each model-task combination. After collect_runs() builds the initial list, a filtering step (lines 16-23 in evaluation/compare.py) compares timestamp-named directories and discards older entries, ensuring aggregates reflect the latest evaluation state.

What is the difference between pooled and macro criterion-pass rates?

The pooled criterion-pass rate treats all criteria across all tasks as a single population, calculating passes divided by total criteria. The macro criterion-pass rate first computes the pass fraction for each individual task, then averages these fractions equally, preventing large tasks from dominating the metric.

Where does the final all-pass rate calculation happen?

The all-pass rate emerges in _aggregate_across_tasks() at lines 102-104 of evaluation/compare.py. The function sums all_pass_points from all tasks (accounting for dual-judge fractional agreement), casts the result to an integer for single-judge compatibility, and divides by the number of tasks to produce the final rate used for leaderboard ranking.

Can I generate comparison dashboards for specific practice areas only?

Yes. The compare_area() function accepts an area parameter to filter runs by practice domain before aggregation. This targeted approach evaluates only relevant subsets of the results/ directory, producing focused dashboards without processing unrelated evaluation data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →