How to Interpret Criterion-Level Results in Harvey-Labs `scores.json`

Harvey-Labs writes detailed per-criterion outcomes to results/<run_id>/scores.json, where each criterion's verdict and reasoning fields reveal exactly why a task passed or failed.

The scores.json file is the primary artifact for diagnosing model performance in the harveyai/harvey-labs evaluation framework. It combines aggregate task-level metrics with granular criterion-level results, enabling precise debugging of where and why a legal AI agent succeeded or fell short.

Understanding the scores.json File Structure

When an evaluation run completes, evaluation/run_eval.py writes the scores.json file to results/<run_id>/scores.json【https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py#L56-L60】. The file has two distinct sections: top-level summary fields and the detailed criteria_results array.

Top-Level Aggregate Fields

Field Purpose
run_id Unique identifier combining task, model, and timestamp
score / max_score 1.0 if every criterion passed, otherwise 0.0
all_pass Boolean equivalent of score == 1.0
n_criteria Total rubric criteria evaluated
n_passed Count of criteria with "pass" verdict
summary Human-readable overview (e.g., "8/12 criteria passed. Missed 4 — task FAIL.")

These aggregate fields reflect Harvey-Labs' all-pass grading scheme: a single failed criterion causes the entire task to fail, mirroring real-world legal standards where one missed issue can invalidate a deliverable【https://github.com/harveyai/harvey-labs/blob/main/docs/eval-strategies.md#L92-L99】.

Reading Criterion-Level Results

The criteria_results array contains one entry per rubric criterion. Each element follows this schema:

{
  "id": "C-001",
  "title": "Identifies change-of-control provisions",
  "verdict": "pass",
  "reasoning": "The agent identified all relevant CoC provisions..."
}

Field-by-Field Breakdown

  • id — Internal rubric identifier (e.g., C-001) mapping directly to the task's rubric.json
  • title — Human-readable description of what the criterion evaluates
  • verdict — Binary "pass" or "fail" returned by the LLM judge
  • reasoning — Full textual justification generated by the judge; this is your per-criterion reasoning

The binary verdict originates from Judge._VERDICT_SCHEMA in evaluation/judge.py, which enforces strict pass/fail output【https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py#L20-L27】. The reasoning field captures nuance through the rubric_criterion prompt template in evaluation/prompts/rubric_criterion.txt【https://github.com/harveyai/harvey-labs/blob/main/docs/eval-strategies.md#L20-L23】.

Practical Workflows for Per-Criterion Analysis

Diagnose Failures with Precise Reasoning

When verdict is "fail", the reasoning field explains exactly which match_criteria requirement was unmet. This eliminates guesswork about whether the model missed content, misidentified it, or produced incorrect analysis.

Audit Judge Consistency

Review reasoning strings across all criteria to verify uniform application of interpretive standards. Inconsistent reasoning patterns may indicate prompt instability or rubric ambiguity requiring refinement.

Measure Granularity for Iteration

Use n_criteria and n_passed for quick progress tracking, then drill into criteria_results for targeted model improvements or rubric adjustments.

Loading and Analyzing scores.json Programmatically

Display a Concise Criterion Table

import json
from pathlib import Path

def load_scores(path: Path) -> dict:
    """Read a scores.json file and return the parsed dict."""
    return json.loads(path.read_text())

def print_criterion_table(scores: dict) -> None:
    """Display each criterion's verdict in a markdown-style table."""
    print("| ID | Title | Verdict |")
    print("|----|-------|---------|")
    for c in scores["criteria_results"]:
        print(f"| {c['id']} | {c['title']} | {c['verdict']} |")

# Usage

scores_path = Path(
    "results/real-estate/extract-psa-key-terms/scenario-01/"
    "claude-sonnet-4-6-20260428-142301/scores.json"
)
scores = load_scores(scores_path)
print_criterion_table(scores)

Extract Reasoning for Failed Criteria Only

def failed_reasonings(scores: dict) -> list[tuple[str, str]]:
    """Return (criterion_id, reasoning) tuples for all failed items."""
    return [
        (c["id"], c["reasoning"])
        for c in scores["criteria_results"]
        if c["verdict"] == "fail"
    ]

failed = failed_reasonings(scores)
for cid, reasoning in failed:
    print(f"---\nCriterion {cid} failed because:\n{reasoning}\n")

These patterns support automated reporting pipelines and manual error analysis workflows.

Why Criterion-Level Reasoning Matters

The Harvey-Labs evaluation design intentionally restricts numerical scores to binary outcomes while preserving explanatory depth in text fields. This approach:

  • Prevents false precision — Legal tasks rarely admit partial credit; a missed indemnification clause is a failure regardless of "75% completion"
  • Enables targeted fixes — Detailed reasoning directs engineering effort toward specific model weaknesses rather than aggregate metrics
  • Supports regulatory review — Verbatim judge reasoning provides auditable evidence of evaluation methodology

Summary

  • Location: results/<run_id>/scores.json written by evaluation/run_eval.py
  • Aggregate fields: score, all_pass, n_criteria, n_passed for task-level outcomes
  • Per-criterion data: criteria_results array with id, title, verdict, reasoning
  • Binary verdicts: Enforced by Judge._VERDICT_SCHEMA in evaluation/judge.py with no partial credit
  • Reasoning source: Generated via evaluation/prompts/rubric_criterion.txt prompt template
  • Key workflow: Load JSON → filter by verdict → read reasoning to diagnose failures

Frequently Asked Questions

Where is the scores.json file located after a run?

The file is written to results/<run_id>/scores.json by evaluation/run_eval.py immediately after the judge completes scoring all criteria. The <run_id> incorporates the task name, model identifier, and timestamp.

Can a criterion receive a partial score between pass and fail?

No. Harvey-Labs uses strict binary grading where every criterion must receive either "pass" or "fail". This all-pass scheme reflects legal practice standards where incomplete analysis of a critical issue constitutes failure. The reasoning field captures any nuance about performance quality.

How do I map criterion IDs back to the original rubric?

Each id in criteria_results (e.g., C-001) corresponds directly to the criterion definition in the task's rubric.json. The title field provides a human-readable cross-reference, making manual alignment straightforward.

What determines the content of the reasoning field?

The judge generates reasoning using the rubric_criterion.txt prompt template, which instructs the LLM to explain its verdict with reference to specific match_criteria requirements. The response is stored verbatim in the reasoning field without modification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →