# Single-Judge vs Dual-Judge Evaluation Modes in harvey-labs run_eval.py

> Understand single-judge vs dual-judge evaluation modes in harvey-labs run_eval.py. Learn how each mode uses LLMs to score benchmarks and generates different JSON outputs.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: internals
- Published: 2026-08-11

---

**Single-judge mode grades benchmark runs with one LLM and produces a single [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json), while dual-judge mode employs both standard LAB judges (`claude-sonnet-4-6` and `gpt-5.5`), averages their independent verdicts, and generates per-model reports plus a consolidated [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json).**

The [`run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/run_eval.py) script in the harvey-labs repository provides two distinct strategies for evaluating benchmark runs. Understanding the difference between single-judge and dual-judge evaluation modes allows you to choose between rapid iteration with a single model or robust, bias-mitigated validation using dual consensus.

## How Single-Judge Mode Works

### Invocation and Configuration

Single-judge evaluation is the default behavior when executing [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py). Use the `--judge-model` argument to specify the LLM, which defaults to `claude-sonnet-4-6` if omitted.

```bash
uv run python -m evaluation.run_eval \
    --run-id myrun123 \
    --task real-estate/extract-psa-key-terms/scenario-01 \
    --judge-model gpt-5.5

```

### Execution Flow

In [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) (lines 99-123), the `evaluate_run` function processes the benchmark with a single `Judge` instance. The script instantiates `Judge(model=args.judge_model)` and performs one grading pass, producing a final verdict, numerical score, and summary for that specific model.

### Output Structure

Results are written to `results/<run_id>/scores.json`. The CLI displays a concise summary via the internal `_print_summary` helper, showing metrics derived exclusively from the single judge's assessment.

## How Dual-Judge Mode Works

### Invocation with the --dual Flag

Activate dual-judge evaluation by passing the `--dual` flag. This mode ignores the `--judge-model` argument and automatically uses both standard judges defined in the source code.

```bash
uv run python -m evaluation.run_eval \
    --run-id myrun123 \
    --task real-estate/extract-psa-key-terms/scenario-01 \
    --dual

```

### The Two-Judge Architecture

As defined in [`run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/run_eval.py) (line 29), the `JUDGE_MODELS` constant contains `claude-sonnet-4-6` and `gpt-5.5`. The `evaluate_run_dual` function (around lines 63-85) iterates through these models, creating an independent `Judge` instance for each. Each model executes `evaluate_run` separately, with intermediate results temporarily stored before aggregation.

### Aggregation Logic and Output

After both judges complete, `evaluate_run_dual` calculates averaged pass rates: `dual_crit` (criterion-pass fraction) and `dual_ap` (overall-pass rate). The final output includes three distinct files:

- `results/<run_id>/scores_claude-sonnet-4-6.json`
- `results/<run_id>/scores_gpt-5.5.json`
- `results/<run_id>/scores_dual.json` (containing the averaged metrics)

The CLI displays per-judge breakdowns followed by the dual-averaged summary via `_print_dual_summary`.

## Key Differences at a Glance

| Aspect | Single-Judge Mode | Dual-Judge Mode |
|--------|-------------------|-----------------|
| **CLI Argument** | `--judge-model <model>` | `--dual` |
| **Models Used** | One user-specified LLM (default: `claude-sonnet-4-6`) | Both `claude-sonnet-4-6` and `gpt-5.5` as defined in `JUDGE_MODELS` |
| **Execution** | Single call to `evaluate_run` | Two independent calls to `evaluate_run`, once per model |
| **Scoring Logic** | Direct output from one judge | Averaged criterion-pass and overall-pass rates across both judges |
| **Result Files** | [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) | `scores_<model>.json` for each judge, plus aggregated [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json) |
| **Summary Method** | `_print_summary` (single model) | `_print_dual_summary` (per-judge + averaged) |

## Implementation Details

The evaluation logic resides in [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py). Single-judge flow parses `--judge-model` (lines 67-85) and invokes `evaluate_run` directly. Dual-judge flow executes a loop over `JUDGE_MODELS`, renames temporary [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) files to per-model variants (e.g., [`scores_claude-sonnet-4-6.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_claude-sonnet-4-6.json)), then computes final averages.

Both modes leverage [`evaluation/judge.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py) for LLM provider abstraction and [`evaluation/report.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/report.py) for human-readable report generation.

## Summary

- **Single-judge mode** uses one LLM (specified via `--judge-model`) to generate a single [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) for rapid feedback.
- **Dual-judge mode** runs both standard LAB judges (`claude-sonnet-4-6` and `gpt-5.5`) via the `--dual` flag, mitigating model bias by averaging results into [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json).
- Single-judge relies on `evaluate_run`; dual-judge uses `evaluate_run_dual` with per-model file renaming and metric aggregation.
- Choose single-judge for development speed and dual-judge for production-grade robustness.

## Frequently Asked Questions

### What is the default judge model for single-judge mode?

The default is `claude-sonnet-4-6`. You can override this by passing `--judge-model` followed by your preferred model identifier, such as `gpt-5.5`.

### How does dual-judge mode calculate the final score?

It averages two specific metrics across both judges: the **criterion-pass fraction** (`dual_crit`) and the **overall-pass rate** (`dual_ap`). These values are written to [`scores_dual.json`](https://github.com/harveyai/harvey-labs/blob/main/scores_dual.json), while individual judge scores remain available in their respective per-model files.

### Can I use custom models in dual-judge mode?

No. The `--dual` flag strictly uses the models hardcoded in the `JUDGE_MODELS` constant (`claude-sonnet-4-6` and `gpt-5.5`). To evaluate with custom model pairs, you would need to modify [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) directly.

### Which mode should I use for production evaluations?

Use **dual-judge mode** for production or high-stakes evaluations where mitigating individual model bias is critical. The averaged scores provide more robust validation. Use **single-judge mode** for development iterations where execution speed matters more than cross-model consensus.