Single-Judge vs Dual-Judge Evaluation Modes in harvey-labs run_eval.py

Single-judge mode grades benchmark runs with one LLM and produces a single scores.json, while dual-judge mode employs both standard LAB judges (claude-sonnet-4-6 and gpt-5.5), averages their independent verdicts, and generates per-model reports plus a consolidated scores_dual.json.

The run_eval.py script in the harvey-labs repository provides two distinct strategies for evaluating benchmark runs. Understanding the difference between single-judge and dual-judge evaluation modes allows you to choose between rapid iteration with a single model or robust, bias-mitigated validation using dual consensus.

How Single-Judge Mode Works

Invocation and Configuration

Single-judge evaluation is the default behavior when executing evaluation/run_eval.py. Use the --judge-model argument to specify the LLM, which defaults to claude-sonnet-4-6 if omitted.

uv run python -m evaluation.run_eval \
    --run-id myrun123 \
    --task real-estate/extract-psa-key-terms/scenario-01 \
    --judge-model gpt-5.5

Execution Flow

In evaluation/run_eval.py (lines 99-123), the evaluate_run function processes the benchmark with a single Judge instance. The script instantiates Judge(model=args.judge_model) and performs one grading pass, producing a final verdict, numerical score, and summary for that specific model.

Output Structure

Results are written to results/<run_id>/scores.json. The CLI displays a concise summary via the internal _print_summary helper, showing metrics derived exclusively from the single judge's assessment.

How Dual-Judge Mode Works

Invocation with the --dual Flag

Activate dual-judge evaluation by passing the --dual flag. This mode ignores the --judge-model argument and automatically uses both standard judges defined in the source code.

uv run python -m evaluation.run_eval \
    --run-id myrun123 \
    --task real-estate/extract-psa-key-terms/scenario-01 \
    --dual

The Two-Judge Architecture

As defined in run_eval.py (line 29), the JUDGE_MODELS constant contains claude-sonnet-4-6 and gpt-5.5. The evaluate_run_dual function (around lines 63-85) iterates through these models, creating an independent Judge instance for each. Each model executes evaluate_run separately, with intermediate results temporarily stored before aggregation.

Aggregation Logic and Output

After both judges complete, evaluate_run_dual calculates averaged pass rates: dual_crit (criterion-pass fraction) and dual_ap (overall-pass rate). The final output includes three distinct files:

  • results/<run_id>/scores_claude-sonnet-4-6.json
  • results/<run_id>/scores_gpt-5.5.json
  • results/<run_id>/scores_dual.json (containing the averaged metrics)

The CLI displays per-judge breakdowns followed by the dual-averaged summary via _print_dual_summary.

Key Differences at a Glance

Aspect Single-Judge Mode Dual-Judge Mode
CLI Argument --judge-model <model> --dual
Models Used One user-specified LLM (default: claude-sonnet-4-6) Both claude-sonnet-4-6 and gpt-5.5 as defined in JUDGE_MODELS
Execution Single call to evaluate_run Two independent calls to evaluate_run, once per model
Scoring Logic Direct output from one judge Averaged criterion-pass and overall-pass rates across both judges
Result Files scores.json scores_<model>.json for each judge, plus aggregated scores_dual.json
Summary Method _print_summary (single model) _print_dual_summary (per-judge + averaged)

Implementation Details

The evaluation logic resides in evaluation/run_eval.py. Single-judge flow parses --judge-model (lines 67-85) and invokes evaluate_run directly. Dual-judge flow executes a loop over JUDGE_MODELS, renames temporary scores.json files to per-model variants (e.g., scores_claude-sonnet-4-6.json), then computes final averages.

Both modes leverage evaluation/judge.py for LLM provider abstraction and evaluation/report.py for human-readable report generation.

Summary

  • Single-judge mode uses one LLM (specified via --judge-model) to generate a single scores.json for rapid feedback.
  • Dual-judge mode runs both standard LAB judges (claude-sonnet-4-6 and gpt-5.5) via the --dual flag, mitigating model bias by averaging results into scores_dual.json.
  • Single-judge relies on evaluate_run; dual-judge uses evaluate_run_dual with per-model file renaming and metric aggregation.
  • Choose single-judge for development speed and dual-judge for production-grade robustness.

Frequently Asked Questions

What is the default judge model for single-judge mode?

The default is claude-sonnet-4-6. You can override this by passing --judge-model followed by your preferred model identifier, such as gpt-5.5.

How does dual-judge mode calculate the final score?

It averages two specific metrics across both judges: the criterion-pass fraction (dual_crit) and the overall-pass rate (dual_ap). These values are written to scores_dual.json, while individual judge scores remain available in their respective per-model files.

Can I use custom models in dual-judge mode?

No. The --dual flag strictly uses the models hardcoded in the JUDGE_MODELS constant (claude-sonnet-4-6 and gpt-5.5). To evaluate with custom model pairs, you would need to modify evaluation/run_eval.py directly.

Which mode should I use for production evaluations?

Use dual-judge mode for production or high-stakes evaluations where mitigating individual model bias is critical. The averaged scores provide more robust validation. Use single-judge mode for development iterations where execution speed matters more than cross-model consensus.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →