# Debugging Evaluation Failures in i-have-adhd: A Complete Troubleshooting Guide

> Debug evaluation failures in i-have-adhd with this complete troubleshooting guide. Learn to identify and fix data problems in case files and score rows using the repository's validation pipeline.

- Repository: [Ayoub Ghriss/i-have-adhd](https://github.com/ayghri/i-have-adhd)
- Tags: how-to-guide
- Published: 2026-08-03

---

**The i-have-adhd repository uses a three-stage validation pipeline in [`scripts/run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/scripts/run_evals.py) to catch evaluation failures before running LLM calls, with descriptive error messages pointing to exact data problems in case files or score rows.**

When working with the evaluation harness in [ayghri/i-have-adhd](https://github.com/ayghri/i-have-adhd), debugging evaluation failures requires understanding how the validation pipeline isolates data errors from runtime execution issues. The harness validates case catalogs, score rows, and condition pairings before any LLM calls are made, ensuring you get actionable error messages instead of cryptic runtime failures.

## The Three-Stage Validation Pipeline

The debugging process for evaluation failures centers on three validation stages implemented in [`run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/run_evals.py). Each stage raises specific exceptions that identify exactly what's wrong with your evaluation data.

### Stage 1: Case Catalog Validation

The `validate_cases` function (lines 59-79 in [`scripts/run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/scripts/run_evals.py)) checks every entry in `evals/cases.jsonl` for structural integrity.

This validator ensures:
- All required fields (`id`, `category`, `prompt`, `risk`, `criteria`) are present
- Case IDs are unique across the catalog
- The `risk` field contains only `low`, `medium`, or `high`
- The `criteria` list is non-empty

Typical failure messages include `Case 3: missing fields: prompt, risk` or `Duplicate case id: direct-answer`.

```bash
python -m scripts.run_evals validate --cases evals/cases.jsonl

```

A successful validation prints **"Evaluation cases are valid."**; failures list every problem found with row references.

### Stage 2: Score Row Validation

The `_validate_score` function (lines 82-95) validates submitted scoring data against the expected schema.

Each score row must contain:
- All five metrics: `correctness`, `autonomy`, `actionability`, `safety`, `concision` (values 1-5)
- A `condition` field with value `baseline`, `candidate`, or `comparator`
- Boolean `blocker` flag
- Optional `notes` string

Failures produce messages like `Score row 7: missing fields: autonomy` or `must be between 1 and 5 for correctness`.

### Stage 3: Pairing and Completeness Verification

The `_check_pairing` function (lines 101-128) ensures fair comparison across experimental conditions.

This critical check verifies that:
- Every condition was judged on identical case-trial pairs
- No duplicate score rows exist for the same case-trial-condition combination
- The evaluation matrix is complete

Common failure: `candidate was not judged on the same rows as baseline: missing direct-answer/trial 2; unmatched refusal/trial 1`

## Step-by-Step Debugging Workflow

Follow this sequence to resolve evaluation failures efficiently.

### Step 1: Run the Validator in Isolation

Always start with the `validate` sub-command to catch data errors before consuming API budget:

```bash
python -m scripts.run_evals validate

```

### Step 2: Parse the Exception Traceback

The harness raises `ValueError` with descriptive messages. The traceback points to the exact line in [`run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/run_evals.py), and the message includes the offending row number or case ID.

For example, a duplicate row error:

```

ValueError: Score row 12: duplicate score rows for direct-answer/trial 1

```

### Step 3: Fix the Source Data

| Error Type | Fix Location | Action |
|-----------|-------------|--------|
| Missing/duplicate case IDs | `evals/cases.jsonl` | Add required fields; ensure ID uniqueness |
| Invalid risk values | `evals/cases.jsonl` | Use only `low`, `medium`, `high` |
| Malformed score rows | Your scores JSONL file | Ensure all five metrics present, range 1-5 |
| Pairing mismatch | Your scores JSONL file | Align case-trial pairs across all conditions |

### Step 4: Regenerate the Run Matrix (If Needed)

For pairing failures, use the `plan` command to see the expected evaluation structure:

```bash
python -m scripts.run_evals plan \
  --cases evals/cases.jsonl \
  --trials 2 \
  --include-comparator

```

This outputs the complete matrix of `{case_id, trial, condition}` rows that your scoring data must cover.

### Step 5: Verify Budget Constraints

If evaluation stops with **"Budget exhausted"**, check:
- The `--budget-usd` parameter value
- That your runner correctly reports `cost_usd` in its output

The harness aborts after the call that exceeds budget, showing the final **"Reported cost:"** line for diagnosis.

## Using the Test Suite to Understand Failures

The test file [`tests/test_run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/tests/test_run_evals.py) deliberately triggers each validation error. Run it to see expected failure patterns:

```bash
pytest tests/test_run_evals.py -q

```

Each test case demonstrates the exact message format you'll encounter in production, making it a reliable reference for interpreting real failures.

## Key Files for Debugging Evaluation Failures

| File | Purpose | Critical Functions |
|------|---------|------------------|
| [`scripts/run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/scripts/run_evals.py) | Main evaluation driver | `validate_cases`, `_validate_score`, `_check_pairing` |
| [`tests/test_run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/tests/test_run_evals.py) | Validation test suite | Exercises all error paths |
| `evals/cases.jsonl` | Case catalog schema | Validated by `validate_cases` (L59-79) |
| [`evals/runners.example.json`](https://github.com/ayghri/i-have-adhd/blob/main/evals/runners.example.json) | LLM runner configuration | Referenced by `run` sub-command |

## Summary

- **Running `validate`** isolates data errors before expensive LLM calls
- **Three validation stages** check cases, scores, and pairings with specific error messages
- **Tracebacks include row numbers** and case IDs for precise debugging
- **The `plan` command** generates the expected evaluation matrix for alignment checks
- **Tests in [`test_run_evals.py`](https://github.com/ayghri/i-have-adhd/blob/main/test_run_evals.py)** demonstrate all failure modes and expected fixes

## Frequently Asked Questions

### What does "unsupported condition" mean in an evaluation failure?

This error from `_validate_score` indicates the `condition` field in your score row contains a value other than `baseline`, `candidate`, or `comparator`. Check your scoring data and use only these three valid condition labels.

### How do I fix "candidate was not judged on the same rows as baseline"?

This `_check_pairing` failure means your experimental conditions have mismatched case-trial coverage. Use `python -m scripts.run_evals plan` to generate the complete matrix, then ensure every condition has scores for exactly those case-trial pairs—no more, no less.

### Why does validation pass but the run command still fails?

The `validate` sub-command only checks `evals/cases.jsonl`. Runtime failures typically involve: (1) your runner configuration in [`evals/runners.example.json`](https://github.com/ayghri/i-have-adhd/blob/main/evals/runners.example.json), (2) missing or malformed score files for the `score` sub-command, or (3) pairing validation that only runs when processing actual scores. Run `validate` on your scores file explicitly if using manual scoring.

### Can I debug evaluation failures without running LLM calls?

Yes—use the `validate` and `plan` sub-commands, plus the test suite. These exercise the full validation pipeline without invoking any LLM backends or consuming API budget.