# How to Interpret Criterion-Level Results in Harvey-Labs `scores.json`

> Decode harvey-labs scores.json criterion-level results. Understand verdict and reasoning for task pass or fail to gain deep insights into your test outcomes.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: how-to-guide
- Published: 2026-08-11

---

**Harvey-Labs writes detailed per-criterion outcomes to `results/<run_id>/scores.json`, where each criterion's `verdict` and `reasoning` fields reveal exactly why a task passed or failed.**

The [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) file is the primary artifact for diagnosing model performance in the [harveyai/harvey-labs](https://github.com/harveyai/harvey-labs) evaluation framework. It combines aggregate task-level metrics with granular criterion-level results, enabling precise debugging of where and why a legal AI agent succeeded or fell short.

## Understanding the [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) File Structure

When an evaluation run completes, [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) writes the [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) file to `results/<run_id>/scores.json`【https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py#L56-L60】. The file has two distinct sections: top-level summary fields and the detailed `criteria_results` array.

### Top-Level Aggregate Fields

| Field | Purpose |
|-------|---------|
| `run_id` | Unique identifier combining task, model, and timestamp |
| `score` / `max_score` | `1.0` if **every** criterion passed, otherwise `0.0` |
| `all_pass` | Boolean equivalent of `score == 1.0` |
| `n_criteria` | Total rubric criteria evaluated |
| `n_passed` | Count of criteria with `"pass"` verdict |
| `summary` | Human-readable overview (e.g., "8/12 criteria passed. Missed 4 — task FAIL.") |

These aggregate fields reflect Harvey-Labs' **all-pass grading scheme**: a single failed criterion causes the entire task to fail, mirroring real-world legal standards where one missed issue can invalidate a deliverable【https://github.com/harveyai/harvey-labs/blob/main/docs/eval-strategies.md#L92-L99】.

## Reading Criterion-Level Results

The `criteria_results` array contains one entry per rubric criterion. Each element follows this schema:

```json
{
  "id": "C-001",
  "title": "Identifies change-of-control provisions",
  "verdict": "pass",
  "reasoning": "The agent identified all relevant CoC provisions..."
}

```

### Field-by-Field Breakdown

- **`id`** — Internal rubric identifier (e.g., `C-001`) mapping directly to the task's [`rubric.json`](https://github.com/harveyai/harvey-labs/blob/main/rubric.json)
- **`title`** — Human-readable description of what the criterion evaluates
- **`verdict`** — Binary `"pass"` or `"fail"` returned by the LLM judge
- **`reasoning`** — Full textual justification generated by the judge; this is your **per-criterion reasoning**

The binary verdict originates from `Judge._VERDICT_SCHEMA` in [`evaluation/judge.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py), which enforces strict pass/fail output【https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py#L20-L27】. The `reasoning` field captures nuance through the `rubric_criterion` prompt template in [`evaluation/prompts/rubric_criterion.txt`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/prompts/rubric_criterion.txt)【https://github.com/harveyai/harvey-labs/blob/main/docs/eval-strategies.md#L20-L23】.

## Practical Workflows for Per-Criterion Analysis

### Diagnose Failures with Precise Reasoning

When `verdict` is `"fail"`, the `reasoning` field explains exactly which `match_criteria` requirement was unmet. This eliminates guesswork about whether the model missed content, misidentified it, or produced incorrect analysis.

### Audit Judge Consistency

Review `reasoning` strings across all criteria to verify uniform application of interpretive standards. Inconsistent reasoning patterns may indicate prompt instability or rubric ambiguity requiring refinement.

### Measure Granularity for Iteration

Use `n_criteria` and `n_passed` for quick progress tracking, then drill into `criteria_results` for targeted model improvements or rubric adjustments.

## Loading and Analyzing [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) Programmatically

### Display a Concise Criterion Table

```python
import json
from pathlib import Path

def load_scores(path: Path) -> dict:
    """Read a scores.json file and return the parsed dict."""
    return json.loads(path.read_text())

def print_criterion_table(scores: dict) -> None:
    """Display each criterion's verdict in a markdown-style table."""
    print("| ID | Title | Verdict |")
    print("|----|-------|---------|")
    for c in scores["criteria_results"]:
        print(f"| {c['id']} | {c['title']} | {c['verdict']} |")

# Usage

scores_path = Path(
    "results/real-estate/extract-psa-key-terms/scenario-01/"
    "claude-sonnet-4-6-20260428-142301/scores.json"
)
scores = load_scores(scores_path)
print_criterion_table(scores)

```

### Extract Reasoning for Failed Criteria Only

```python
def failed_reasonings(scores: dict) -> list[tuple[str, str]]:
    """Return (criterion_id, reasoning) tuples for all failed items."""
    return [
        (c["id"], c["reasoning"])
        for c in scores["criteria_results"]
        if c["verdict"] == "fail"
    ]

failed = failed_reasonings(scores)
for cid, reasoning in failed:
    print(f"---\nCriterion {cid} failed because:\n{reasoning}\n")

```

These patterns support automated reporting pipelines and manual error analysis workflows.

## Why Criterion-Level Reasoning Matters

The Harvey-Labs evaluation design intentionally restricts numerical scores to binary outcomes while preserving explanatory depth in text fields. This approach:

- **Prevents false precision** — Legal tasks rarely admit partial credit; a missed indemnification clause is a failure regardless of "75% completion"
- **Enables targeted fixes** — Detailed reasoning directs engineering effort toward specific model weaknesses rather than aggregate metrics
- **Supports regulatory review** — Verbatim judge reasoning provides auditable evidence of evaluation methodology

## Summary

- **Location**: `results/<run_id>/scores.json` written by [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py)
- **Aggregate fields**: `score`, `all_pass`, `n_criteria`, `n_passed` for task-level outcomes
- **Per-criterion data**: `criteria_results` array with `id`, `title`, `verdict`, `reasoning`
- **Binary verdicts**: Enforced by `Judge._VERDICT_SCHEMA` in [`evaluation/judge.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py) with no partial credit
- **Reasoning source**: Generated via [`evaluation/prompts/rubric_criterion.txt`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/prompts/rubric_criterion.txt) prompt template
- **Key workflow**: Load JSON → filter by `verdict` → read `reasoning` to diagnose failures

## Frequently Asked Questions

### Where is the [`scores.json`](https://github.com/harveyai/harvey-labs/blob/main/scores.json) file located after a run?

The file is written to `results/<run_id>/scores.json` by [`evaluation/run_eval.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/run_eval.py) immediately after the judge completes scoring all criteria. The `<run_id>` incorporates the task name, model identifier, and timestamp.

### Can a criterion receive a partial score between pass and fail?

No. Harvey-Labs uses strict binary grading where every criterion must receive either `"pass"` or `"fail"`. This all-pass scheme reflects legal practice standards where incomplete analysis of a critical issue constitutes failure. The `reasoning` field captures any nuance about performance quality.

### How do I map criterion IDs back to the original rubric?

Each `id` in `criteria_results` (e.g., `C-001`) corresponds directly to the criterion definition in the task's [`rubric.json`](https://github.com/harveyai/harvey-labs/blob/main/rubric.json). The `title` field provides a human-readable cross-reference, making manual alignment straightforward.

### What determines the content of the `reasoning` field?

The judge generates reasoning using the [`rubric_criterion.txt`](https://github.com/harveyai/harvey-labs/blob/main/rubric_criterion.txt) prompt template, which instructs the LLM to explain its verdict with reference to specific `match_criteria` requirements. The response is stored verbatim in the `reasoning` field without modification.