# How All-Pass Rubric Scoring Works in Harvey-Labs: A Complete Technical Guide

> Discover how all-pass rubric scoring in Harvey-Labs works. This guide details the LLM judge's binary pass/fail evaluation for task criteria and its pipeline propagation.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: deep-dive
- Published: 2026-08-11

---

**All-pass rubric scoring in Harvey-Labs evaluates whether every criterion in a task's rubric receives a "pass" verdict from an LLM judge, producing a binary flag that propagates through single-judge, dual-judge, and aggregation pipelines.**

The **all-pass rubric scoring** mechanism is a core evaluation primitive in the Harvey-Labs benchmark framework. It transforms per-criterion LLM judgments into a single binary indicator of task success, enabling researchers to quickly identify runs that fully satisfy task requirements. This article explains the implementation details across the scoring, evaluation, and reporting stack.

## Per-Criterion Grading and Verdict Collection

The scoring pipeline begins in [`evaluation/scoring.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/scoring.py), where the `score_rubric` function processes each criterion defined in a task's [`task.json`](https://github.com/harveyai/harvey-labs/blob/main/task.json).

For every criterion, the system:

1. Loads the relevant deliverable files
2. Prompts an LLM judge to render a verdict (`"pass"` or `"fail"`)
3. Stores the result in a `CriterionResult` object

The per-criterion results are later flattened into plain dictionaries for downstream processing. This granular approach ensures that the all-pass determination has full visibility into every individual judgment.

## Computing the All-Pass Flag in Single-Judge Mode

After rubric scoring completes, `evaluate_run` in [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) (lines 16-19) computes the all-pass flag with straightforward logic:

```python
n_criteria = len(result.criteria_results)
n_passed   = sum(1 for c in result.criteria_results if c["verdict"] == "pass")
all_pass  = n_criteria > 0 and n_passed == n_criteria

```

The `all_pass` boolean is **strict**: it is `True` only when every rubric criterion is marked `"pass"`. A single failure invalidates the flag. This value is inserted into the run's scores dictionary as `scores["all_pass"]` and persisted to [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json).

## Dual-Judge Mode and Consensus Requirements

Harvey-Labs supports `--dual` evaluation for increased reliability. In this mode, `evaluate_run_dual` in [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) (lines 7-20) aggregates independent judgments from two LLM judges:

```python
dual_ap = sum(1.0 if scores.get("all_pass") else 0.0 for scores in per_judge.values()) / len(per_judge)
all_pass = dual_ap == 1.0  # only True if **both** judges gave all-pass

```

The **consensus rule** is strict: `all_pass` is `True` only when both judges independently award all-pass status. The system also exposes `dual_all_pass_rate` (the fraction of judges giving all-pass) for sensitivity analysis.

## Score Normalization and Aggregation

The `_comparison_scores` helper in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py) (lines 11-25) normalizes scores across single-judge and dual-judge runs. It exposes:

- `all_pass`: the binary flag
- `all_pass_score`: a numeric equivalent (1.0 for all-pass, 0.0 otherwise)

These fields enable consistent aggregation across heterogeneous evaluation configurations. Researchers can compute all-pass rates across model families, task categories, or experimental conditions using these normalized outputs.

## Visualization and Reporting

The all-pass metric surfaces throughout Harvey-Labs reporting infrastructure:

- **Charts**: [`evaluation/charts.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/charts.py) (lines 587-618) generates the *All-Pass distribution* visualization, showing how frequently runs achieve perfect rubric satisfaction
- **HTML Reports**: [`evaluation/report.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/report.py) renders an "ALL PASS" badge when the flag is true, providing immediate visual recognition of successful runs
- **Export Tables**: CSV and JSON outputs include `all_pass` and `all_pass_score` for downstream statistical analysis

## Practical Usage Examples

### Single-Judge Evaluation

```python
from harvey_labs.evaluation.run_eval import evaluate_run
from harvey_labs.evaluation.judge import Judge

judge = Judge(model="claude-sonnet-4-6")
scores = evaluate_run(
    run_id="run-001",
    task="real-estate/extract-psa-key-terms/scenario-01",
    judge=judge,
    parallel=4,
)

print(scores["all_pass"])   # True ⇔ every rubric criterion passed

print(scores["summary"])    # Human-readable summary with ALL-PASS notice

```

### Dual-Judge Evaluation

```python
from harvey_labs.evaluation.run_eval import evaluate_run_dual

aggregate = evaluate_run_dual(
    run_id="run-002",
    task="legal-contracts/review/scenario-03",
    parallel=6,
)

print(aggregate["all_pass"])              # True only if *both* judges gave all-pass

print(aggregate["dual_all_pass_rate"])   # Fraction of judges that gave all-pass (0-1)

```

### Batch Analysis of Aggregated Results

```python
from harvey_labs.evalivation.compare import collect_runs

runs = collect_runs(task_filter="legal-contracts/review")
for r in runs:
    print(r["run_id"], r["all_pass"], r["all_pass_score"])

```

## Key Implementation Files

| File | Role |
|------|------|
| [`evaluation/scoring.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/scoring.py) | Core rubric scoring (`score_rubric`) and `CriterionResult` handling |
| [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) | Orchestrates scoring, computes `all_pass`, writes [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) |
| [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py) | Normalizes scores, exposes `all_pass` & `all_pass_score` |
| [`evaluation/charts.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/charts.py) | Visualizes all-pass distribution across runs and models |
| [`evaluation/report.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/report.py) | Generates HTML reports with ALL PASS badge highlighting |
| [`tests/test_scoring.py`](https://github.com/harveyai/harvey-labs/blob/main/tests/test_scoring.py) | Unit tests verifying all-pass logic (e.g., `test_all_fail_rubric`) |
| [`tests/test_eval_strategies.py`](https://github.com/harveyai/harvey-labs/blob/main/tests/test_eval_strategies.py) | Integration tests for all-pass rubric evaluation behavior |

## Summary

- **All-pass rubric scoring** produces a binary flag indicating whether every criterion in a task rubric received a `"pass"` verdict from the LLM judge.
- The flag is computed in `evaluate_run` by comparing the count of passed criteria against the total criteria count.
- **Dual-judge mode** requires consensus: both judges must independently award all-pass for the aggregate flag to be true.
- Normalized `all_pass_score` values enable cross-run aggregation and statistical analysis.
- The metric propagates through charts, reports, and export tables as a primary indicator of task completion.

## Frequently Asked Questions

### What happens if a rubric has zero criteria?

The all-pass logic in `evaluate_run` explicitly checks `n_criteria > 0`. If no criteria exist, `all_pass` evaluates to `False` regardless of passed count, preventing vacuous truth cases from polluting results.

### Can I use all-pass scoring with custom judge configurations?

Yes. The `all_pass` computation is judge-agnostic—it operates on the verdict strings in `c["verdict"]`. Any judge implementation that returns `"pass"` or `"fail"` per criterion will integrate correctly with the existing pipeline.

### How does dual-judge mode handle judge disagreement on individual criteria?

`evaluate_run_dual` computes `all_pass` per judge independently first, then requires both to be `True` for consensus. The system does not expose per-criterion disagreement directly in the aggregate; use per-judge [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) outputs for granular analysis.

### Where is the all-pass flag persisted for long-term analysis?

Each run's [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) contains the `"all_pass"` boolean. The `collect_runs` function in [`evaluation/compare.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/compare.py) aggregates these across runs, and [`evaluation/charts.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/charts.py) generates distribution visualizations from the collected data.