# How Harvey-Labs Comparison Dashboards Generate Aggregate Metrics Across Multiple Runs

> Discover how Harvey-Labs comparison dashboards aggregate metrics across multiple runs by scanning results, deduplicating, grouping by model, and computing pooled pass rates.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: internals
- Published: 2026-08-11

---

**Harvey-Labs aggregates evaluation runs by scanning the `results/` directory for score files, deduplicating by timestamp, grouping runs by model label, and computing pooled and macro pass rates to power its comparison dashboards.**

The harvey-labs repository provides a comprehensive evaluation framework for legal AI models, with comparison dashboards that synthesize results from dozens or hundreds of individual test runs. Understanding how harvey-labs comparison dashboards generate aggregate metrics across multiple runs reveals the pipeline that transforms raw JSON scores into leaderboard-ready statistics.

## Scanning and Normalizing Individual Runs

The aggregation process begins in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py), where the `collect_runs()` function (lines 52-68) walks the `results/` directory tree to discover evaluation artifacts. It identifies two distinct file types: [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) for single-judge evaluations and [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json) for dual-judge setups.

Each discovered file passes through `_comparison_scores()`, which standardizes the raw JSON into a common schema. For dual-judge runs, this helper delegates to `_normalize_dual_scores()` imported from [`evaluation/report.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/report.py) (lines 11-30). This normalization ensures that fields like **model id**, **effort**, **token usage**, **cost**, and **pass/fail counts** are uniformly accessible regardless of the evaluation mode. The extracted data accumulates in a temporary list named `raw_runs` before moving to the next stage.

## Deduplicating Repeated Evaluations

When a model undergoes multiple evaluations against the same task, the system retains only the newest run to prevent skewed aggregates. This deduplication logic (lines 16-23 in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py)) compares timestamp-named directories and filters the `raw_runs` list to keep the latest entry per model-task combination.

## Aggregating Metrics Across Tasks

The core mathematical aggregation happens in `_aggregate_across_tasks()` (lines 26-41), which collapses the flat list of runs into per-model summary statistics.

### Grouping by Model Label

Runs cluster by their human-readable `pretty_label`, which encodes the model identifier and effort level. This grouping ensures that variants of the same model are analyzed separately while all tasks for a specific configuration roll up together.

### Computing Pass Rates and Costs

For each model group, the function tallies cumulative values:
- `total_passed` and `total_criteria` for raw accuracy counts
- `total_tokens`, `total_wall_clock`, and `total_cost` for resource consumption
- `criterion_pass_fraction_sum` for per-task pass ratios
- `all_pass_points` for dual-judge agreement scores (which may be fractional)

From these sums, the system derives three critical metrics. The **pooled criterion-pass rate** (lines 84-89) divides total passes by total criteria across all tasks. The **macro criterion-pass rate** (lines 90-94) averages the per-task pass fractions, giving equal weight to each task regardless of its criteria count.

### Deriving the All-Pass Score

The final leaderboard metric, the **all-pass rate**, originates from the summed `all_pass_points` (lines 96-101). For legacy single-judge runs, this value casts to an integer, while dual-judge runs maintain fractional agreement weights. The rate equals `all_pass_count / n_tasks` (lines 102-104). The function also tracks `all_pass_both_agree_*` counters to record judge concordance.

After calculation, the resulting dictionaries sort by the overall `score` (the all-pass rate) before returning to the caller (lines 138-140).

## Rendering the Dashboard Visualizations

With aggregated data prepared, the `compare_area()` and `compare_all()` commands orchestrate visualization. These functions invoke the collection and aggregation pipeline, then hand the processed dictionaries to charting helpers in [`evaluation/charts.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/charts.py), including `leaderboard_table()`, `grouped_bars()`, and `rubric_vs_allpass_bars()`.

Matplotlib figures generate as either inline displays or PNG files. The `_write_html()` function (lines 31-38) embeds these assets into static HTML reports, producing the final comparison dashboards stored in `results/comparisons/`.

## Practical Usage Examples

Generate a complete dashboard for a specific practice area using the high-level API:

```python
from evaluation.compare import compare_area

# Produce HTML report for funds asset management evaluations

compare_area(area="funds-asset-management", save_images=False)

```

For custom reporting or programmatic analysis, access the aggregation pipeline directly:

```python
from evaluation.compare import collect_runs, _aggregate_across_tasks

# Gather all runs from results/

runs = collect_runs()

# Define task scope

task_list = sorted({run["task"] for run in runs})

# Generate aggregates

aggregated = _aggregate_across_tasks(runs=runs, task_list=task_list)

# Access metrics for the top-performing model

top_model = aggregated[0]
print(f"All-pass rate: {top_model['all_pass_rate']}")
print(f"Pooled criterion rate: {top_model['criterion_pass_rate_pooled']}")
print(f"Total cost: ${top_model['total_cost']:.2f}")

```

## Summary

- **Collection**: `collect_runs()` in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py) scans for [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) and [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json) files, normalizing dual-judge outputs via [`evaluation/report.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/report.py).
- **Deduplication**: Only the newest run per model-task pair survives, determined by timestamp directory naming.
- **Aggregation**: `_aggregate_across_tasks()` groups by `pretty_label` and computes pooled rates (total passes / total criteria), macro rates (average of per-task fractions), and the all-pass rate (primary leaderboard metric).
- **Visualization**: [`evaluation/charts.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/charts.py) consumes aggregated dictionaries to render tables and charts, with HTML generation handled by `_write_html()`.
- **Entry Points**: Use `compare_area()` for specific domains or `compare_all()` for comprehensive benchmarks.

## Frequently Asked Questions

### How does harvey-labs handle duplicate runs of the same model?

The system deduplicates by retaining only the newest run for each model-task combination. After `collect_runs()` builds the initial list, a filtering step (lines 16-23 in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py)) compares timestamp-named directories and discards older entries, ensuring aggregates reflect the latest evaluation state.

### What is the difference between pooled and macro criterion-pass rates?

The **pooled criterion-pass rate** treats all criteria across all tasks as a single population, calculating passes divided by total criteria. The **macro criterion-pass rate** first computes the pass fraction for each individual task, then averages these fractions equally, preventing large tasks from dominating the metric.

### Where does the final all-pass rate calculation happen?

The all-pass rate emerges in `_aggregate_across_tasks()` at lines 102-104 of [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py). The function sums `all_pass_points` from all tasks (accounting for dual-judge fractional agreement), casts the result to an integer for single-judge compatibility, and divides by the number of tasks to produce the final rate used for leaderboard ranking.

### Can I generate comparison dashboards for specific practice areas only?

Yes. The `compare_area()` function accepts an `area` parameter to filter runs by practice domain before aggregation. This targeted approach evaluates only relevant subsets of the `results/` directory, producing focused dashboards without processing unrelated evaluation data.