Where to Find the Scoring Rubric for i‑have‑adhd Evaluations
The official scoring rubric for i‑have‑adhd evaluations is located at evals/rubric.md in the repository, defining five weighted scoring dimensions and release‑gate criteria for model assessment.
This guide walks you through locating, understanding, and applying the scoring rubric used to evaluate response quality in the ayghri/i-have-adhd project. The rubric serves as the authoritative source for judging whether a candidate model meets deployment standards.
Locating the Rubric File
The primary scoring rubric lives at a single, canonical path in the repository:
evals/rubric.md— Defines the five weighted dimensions (Correctness, Autonomy, Actionability, Safety, Concision), the 1‑5 scoring scale, and the release‑gate criteria that determine pass/fail status for candidate models.
According to the source code, this file is the single source of truth for evaluation standards. Both the evaluation harness and documentation point to it directly.
How the Rubric Is Used in Practice
The scoring system integrates into the evaluation workflow through three distinct phases.
Validation, Execution, and Scoring
The scripts/run_evals.py harness orchestrates the entire pipeline:
# 1. Validate the case catalog (optional but recommended)
python3 scripts/run_evals.py validate
# 2. Run evaluations for a condition (baseline or candidate)
python3 scripts/run_evals.py run \
--runner claude \
--condition baseline \
--trials 3 \
--budget-usd 12.5 \
--output evals/results/responses.jsonl
# 3. Score manually-judged results using the rubric
python3 scripts/run_evals.py score evals/results/scores.jsonl
The score command reads result rows and aggregates them according to the rubric in evals/rubric.md, producing weighted scores, blocking findings, and a final release‑gate verdict.
Key References to the Rubric
Multiple files in the codebase depend on evals/rubric.md:
| File | Purpose |
|---|---|
evals/rubric.md |
Defines scoring dimensions, weights, and release‑gate criteria |
scripts/run_evals.py |
Evaluation harness that validates cases, runs models, and applies rubric scoring |
evals/README.md |
Documentation covering validation, planning, running, and scoring workflows |
The evals/README.md specifically explains how to validate evaluation cases and apply the scoring rubric during assessments.
Understanding the Scoring Dimensions
The evals/rubric.md file establishes five weighted dimensions for evaluation:
- Correctness — Factual accuracy and logical soundness
- Autonomy — Appropriate level of independent action without unnecessary user prompting
- Actionability — Clear, concrete steps the user can take
- Safety — Avoidance of harmful, misleading, or risky guidance
- Concision — Efficient communication without excess verbosity
Each dimension uses a 1‑5 scale, with weighted aggregation determining final scores and release‑gate status.
Summary
- The scoring rubric for i‑have‑adhd evaluations is located at
evals/rubric.md. - The rubric defines five weighted dimensions and release‑gate criteria for model assessment.
scripts/run_evals.pyimplements the full evaluation pipeline, including rubric‑based scoring.evals/README.mdprovides documentation on running evaluations and applying the rubric.
Frequently Asked Questions
What file contains the official scoring rubric?
The official scoring rubric is stored at evals/rubric.md in the repository root. This markdown file defines all scoring dimensions, weights, and criteria for determining whether a model passes evaluation.
What are the five scoring dimensions in the rubric?
The rubric specifies Correctness, Autonomy, Actionability, Safety, and Concision. Each dimension is scored on a 1‑5 scale and weighted to produce aggregated results that feed into the release‑gate decision.
How do I apply the rubric to evaluation results?
Run python3 scripts/run_evals.py score evals/results/scores.jsonl after completing manual judgment. The command reads scored rows and computes weighted aggregates according to evals/rubric.md, including blocking findings and final pass/fail status.
Where can I find documentation on the full evaluation workflow?
Consult evals/README.md for complete guidance on validating cases, planning evaluation runs, executing tests, and applying the scoring rubric to results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →