How All-Pass Rubric Scoring Works in Harvey-Labs: A Complete Technical Guide

All-pass rubric scoring in Harvey-Labs evaluates whether every criterion in a task's rubric receives a "pass" verdict from an LLM judge, producing a binary flag that propagates through single-judge, dual-judge, and aggregation pipelines.

The all-pass rubric scoring mechanism is a core evaluation primitive in the Harvey-Labs benchmark framework. It transforms per-criterion LLM judgments into a single binary indicator of task success, enabling researchers to quickly identify runs that fully satisfy task requirements. This article explains the implementation details across the scoring, evaluation, and reporting stack.

Per-Criterion Grading and Verdict Collection

The scoring pipeline begins in evaluation/scoring.py, where the score_rubric function processes each criterion defined in a task's task.json.

For every criterion, the system:

  1. Loads the relevant deliverable files
  2. Prompts an LLM judge to render a verdict ("pass" or "fail")
  3. Stores the result in a CriterionResult object

The per-criterion results are later flattened into plain dictionaries for downstream processing. This granular approach ensures that the all-pass determination has full visibility into every individual judgment.

Computing the All-Pass Flag in Single-Judge Mode

After rubric scoring completes, evaluate_run in evaluation/run_eval.py (lines 16-19) computes the all-pass flag with straightforward logic:

n_criteria = len(result.criteria_results)
n_passed   = sum(1 for c in result.criteria_results if c["verdict"] == "pass")
all_pass  = n_criteria > 0 and n_passed == n_criteria

The all_pass boolean is strict: it is True only when every rubric criterion is marked "pass". A single failure invalidates the flag. This value is inserted into the run's scores dictionary as scores["all_pass"] and persisted to scores.json.

Dual-Judge Mode and Consensus Requirements

Harvey-Labs supports --dual evaluation for increased reliability. In this mode, evaluate_run_dual in evaluation/run_eval.py (lines 7-20) aggregates independent judgments from two LLM judges:

dual_ap = sum(1.0 if scores.get("all_pass") else 0.0 for scores in per_judge.values()) / len(per_judge)
all_pass = dual_ap == 1.0  # only True if **both** judges gave all-pass

The consensus rule is strict: all_pass is True only when both judges independently award all-pass status. The system also exposes dual_all_pass_rate (the fraction of judges giving all-pass) for sensitivity analysis.

Score Normalization and Aggregation

The _comparison_scores helper in evaluation/compare.py (lines 11-25) normalizes scores across single-judge and dual-judge runs. It exposes:

  • all_pass: the binary flag
  • all_pass_score: a numeric equivalent (1.0 for all-pass, 0.0 otherwise)

These fields enable consistent aggregation across heterogeneous evaluation configurations. Researchers can compute all-pass rates across model families, task categories, or experimental conditions using these normalized outputs.

Visualization and Reporting

The all-pass metric surfaces throughout Harvey-Labs reporting infrastructure:

  • Charts: evaluation/charts.py (lines 587-618) generates the All-Pass distribution visualization, showing how frequently runs achieve perfect rubric satisfaction
  • HTML Reports: evaluation/report.py renders an "ALL PASS" badge when the flag is true, providing immediate visual recognition of successful runs
  • Export Tables: CSV and JSON outputs include all_pass and all_pass_score for downstream statistical analysis

Practical Usage Examples

Single-Judge Evaluation

from harvey_labs.evaluation.run_eval import evaluate_run
from harvey_labs.evaluation.judge import Judge

judge = Judge(model="claude-sonnet-4-6")
scores = evaluate_run(
    run_id="run-001",
    task="real-estate/extract-psa-key-terms/scenario-01",
    judge=judge,
    parallel=4,
)

print(scores["all_pass"])   # True ⇔ every rubric criterion passed

print(scores["summary"])    # Human-readable summary with ALL-PASS notice

Dual-Judge Evaluation

from harvey_labs.evaluation.run_eval import evaluate_run_dual

aggregate = evaluate_run_dual(
    run_id="run-002",
    task="legal-contracts/review/scenario-03",
    parallel=6,
)

print(aggregate["all_pass"])              # True only if *both* judges gave all-pass

print(aggregate["dual_all_pass_rate"])   # Fraction of judges that gave all-pass (0-1)

Batch Analysis of Aggregated Results

from harvey_labs.evalivation.compare import collect_runs

runs = collect_runs(task_filter="legal-contracts/review")
for r in runs:
    print(r["run_id"], r["all_pass"], r["all_pass_score"])

Key Implementation Files

File Role
evaluation/scoring.py Core rubric scoring (score_rubric) and CriterionResult handling
evaluation/run_eval.py Orchestrates scoring, computes all_pass, writes scores.json
evaluation/compare.py Normalizes scores, exposes all_pass & all_pass_score
evaluation/charts.py Visualizes all-pass distribution across runs and models
evaluation/report.py Generates HTML reports with ALL PASS badge highlighting
tests/test_scoring.py Unit tests verifying all-pass logic (e.g., test_all_fail_rubric)
tests/test_eval_strategies.py Integration tests for all-pass rubric evaluation behavior

Summary

  • All-pass rubric scoring produces a binary flag indicating whether every criterion in a task rubric received a "pass" verdict from the LLM judge.
  • The flag is computed in evaluate_run by comparing the count of passed criteria against the total criteria count.
  • Dual-judge mode requires consensus: both judges must independently award all-pass for the aggregate flag to be true.
  • Normalized all_pass_score values enable cross-run aggregation and statistical analysis.
  • The metric propagates through charts, reports, and export tables as a primary indicator of task completion.

Frequently Asked Questions

What happens if a rubric has zero criteria?

The all-pass logic in evaluate_run explicitly checks n_criteria > 0. If no criteria exist, all_pass evaluates to False regardless of passed count, preventing vacuous truth cases from polluting results.

Can I use all-pass scoring with custom judge configurations?

Yes. The all_pass computation is judge-agnostic—it operates on the verdict strings in c["verdict"]. Any judge implementation that returns "pass" or "fail" per criterion will integrate correctly with the existing pipeline.

How does dual-judge mode handle judge disagreement on individual criteria?

evaluate_run_dual computes all_pass per judge independently first, then requires both to be True for consensus. The system does not expose per-criterion disagreement directly in the aggregate; use per-judge scores.json outputs for granular analysis.

Where is the all-pass flag persisted for long-term analysis?

Each run's scores.json contains the "all_pass" boolean. The collect_runs function in evaluation/compare.py aggregates these across runs, and evaluation/charts.py generates distribution visualizations from the collected data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →