How All-Pass Rubric Scoring Works in Harvey-Labs: A Complete Technical Guide
All-pass rubric scoring in Harvey-Labs evaluates whether every criterion in a task's rubric receives a "pass" verdict from an LLM judge, producing a binary flag that propagates through single-judge, dual-judge, and aggregation pipelines.
The all-pass rubric scoring mechanism is a core evaluation primitive in the Harvey-Labs benchmark framework. It transforms per-criterion LLM judgments into a single binary indicator of task success, enabling researchers to quickly identify runs that fully satisfy task requirements. This article explains the implementation details across the scoring, evaluation, and reporting stack.
Per-Criterion Grading and Verdict Collection
The scoring pipeline begins in evaluation/scoring.py, where the score_rubric function processes each criterion defined in a task's task.json.
For every criterion, the system:
- Loads the relevant deliverable files
- Prompts an LLM judge to render a verdict (
"pass"or"fail") - Stores the result in a
CriterionResultobject
The per-criterion results are later flattened into plain dictionaries for downstream processing. This granular approach ensures that the all-pass determination has full visibility into every individual judgment.
Computing the All-Pass Flag in Single-Judge Mode
After rubric scoring completes, evaluate_run in evaluation/run_eval.py (lines 16-19) computes the all-pass flag with straightforward logic:
n_criteria = len(result.criteria_results)
n_passed = sum(1 for c in result.criteria_results if c["verdict"] == "pass")
all_pass = n_criteria > 0 and n_passed == n_criteria
The all_pass boolean is strict: it is True only when every rubric criterion is marked "pass". A single failure invalidates the flag. This value is inserted into the run's scores dictionary as scores["all_pass"] and persisted to scores.json.
Dual-Judge Mode and Consensus Requirements
Harvey-Labs supports --dual evaluation for increased reliability. In this mode, evaluate_run_dual in evaluation/run_eval.py (lines 7-20) aggregates independent judgments from two LLM judges:
dual_ap = sum(1.0 if scores.get("all_pass") else 0.0 for scores in per_judge.values()) / len(per_judge)
all_pass = dual_ap == 1.0 # only True if **both** judges gave all-pass
The consensus rule is strict: all_pass is True only when both judges independently award all-pass status. The system also exposes dual_all_pass_rate (the fraction of judges giving all-pass) for sensitivity analysis.
Score Normalization and Aggregation
The _comparison_scores helper in evaluation/compare.py (lines 11-25) normalizes scores across single-judge and dual-judge runs. It exposes:
all_pass: the binary flagall_pass_score: a numeric equivalent (1.0 for all-pass, 0.0 otherwise)
These fields enable consistent aggregation across heterogeneous evaluation configurations. Researchers can compute all-pass rates across model families, task categories, or experimental conditions using these normalized outputs.
Visualization and Reporting
The all-pass metric surfaces throughout Harvey-Labs reporting infrastructure:
- Charts:
evaluation/charts.py(lines 587-618) generates the All-Pass distribution visualization, showing how frequently runs achieve perfect rubric satisfaction - HTML Reports:
evaluation/report.pyrenders an "ALL PASS" badge when the flag is true, providing immediate visual recognition of successful runs - Export Tables: CSV and JSON outputs include
all_passandall_pass_scorefor downstream statistical analysis
Practical Usage Examples
Single-Judge Evaluation
from harvey_labs.evaluation.run_eval import evaluate_run
from harvey_labs.evaluation.judge import Judge
judge = Judge(model="claude-sonnet-4-6")
scores = evaluate_run(
run_id="run-001",
task="real-estate/extract-psa-key-terms/scenario-01",
judge=judge,
parallel=4,
)
print(scores["all_pass"]) # True ⇔ every rubric criterion passed
print(scores["summary"]) # Human-readable summary with ALL-PASS notice
Dual-Judge Evaluation
from harvey_labs.evaluation.run_eval import evaluate_run_dual
aggregate = evaluate_run_dual(
run_id="run-002",
task="legal-contracts/review/scenario-03",
parallel=6,
)
print(aggregate["all_pass"]) # True only if *both* judges gave all-pass
print(aggregate["dual_all_pass_rate"]) # Fraction of judges that gave all-pass (0-1)
Batch Analysis of Aggregated Results
from harvey_labs.evalivation.compare import collect_runs
runs = collect_runs(task_filter="legal-contracts/review")
for r in runs:
print(r["run_id"], r["all_pass"], r["all_pass_score"])
Key Implementation Files
| File | Role |
|---|---|
evaluation/scoring.py |
Core rubric scoring (score_rubric) and CriterionResult handling |
evaluation/run_eval.py |
Orchestrates scoring, computes all_pass, writes scores.json |
evaluation/compare.py |
Normalizes scores, exposes all_pass & all_pass_score |
evaluation/charts.py |
Visualizes all-pass distribution across runs and models |
evaluation/report.py |
Generates HTML reports with ALL PASS badge highlighting |
tests/test_scoring.py |
Unit tests verifying all-pass logic (e.g., test_all_fail_rubric) |
tests/test_eval_strategies.py |
Integration tests for all-pass rubric evaluation behavior |
Summary
- All-pass rubric scoring produces a binary flag indicating whether every criterion in a task rubric received a
"pass"verdict from the LLM judge. - The flag is computed in
evaluate_runby comparing the count of passed criteria against the total criteria count. - Dual-judge mode requires consensus: both judges must independently award all-pass for the aggregate flag to be true.
- Normalized
all_pass_scorevalues enable cross-run aggregation and statistical analysis. - The metric propagates through charts, reports, and export tables as a primary indicator of task completion.
Frequently Asked Questions
What happens if a rubric has zero criteria?
The all-pass logic in evaluate_run explicitly checks n_criteria > 0. If no criteria exist, all_pass evaluates to False regardless of passed count, preventing vacuous truth cases from polluting results.
Can I use all-pass scoring with custom judge configurations?
Yes. The all_pass computation is judge-agnostic—it operates on the verdict strings in c["verdict"]. Any judge implementation that returns "pass" or "fail" per criterion will integrate correctly with the existing pipeline.
How does dual-judge mode handle judge disagreement on individual criteria?
evaluate_run_dual computes all_pass per judge independently first, then requires both to be True for consensus. The system does not expose per-criterion disagreement directly in the aggregate; use per-judge scores.json outputs for granular analysis.
Where is the all-pass flag persisted for long-term analysis?
Each run's scores.json contains the "all_pass" boolean. The collect_runs function in evaluation/compare.py aggregates these across runs, and evaluation/charts.py generates distribution visualizations from the collected data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →