Evals Cases JSONL Structure in i-have-adhd: Complete Field Reference
The evals/cases.jsonl file in the i-have-adhd repository uses JSON Lines format where each line is an independent JSON object containing id, category, prompt, risk, and criteria fields.
This file powers the evaluation pipeline for the ayghri/i-have-adhd project, storing test scenarios that validate model behavior across different risk levels and interaction types. Understanding the evals cases JSONL structure is essential for anyone extending the evaluation suite or debugging test failures.
JSON Lines Format Overview
The evals/cases.jsonl file follows the JSON Lines specification—one valid JSON object per line, no trailing commas, no outer array wrapper. This design enables memory-efficient streaming: you can process million-case datasets without loading everything into RAM.
import json
with open('evals/cases.jsonl', 'r') as f:
for line in f:
case = json.loads(line)
# Process each case independently
Core Fields in Each Case Object
Every line in evals/cases.jsonl contains a standardized case object with five required properties:
| Field | Type | Purpose |
|---|---|---|
id |
string | Unique identifier (e.g., "direct-answer", "debug-task-loop") |
category |
string | Classification of interaction type |
prompt |
string | The actual user prompt sent to the model |
risk |
string | Safety priority level: "low", "medium", or "high" |
criteria |
array of strings | Pass/fail conditions for the response |
Field-by-Field Breakdown
id: Unique Case Identifier
The id field provides a stable reference for tracking test results across runs.
case = {"id": "direct-answer", ...}
Use this field to correlate failures in logs with specific scenario definitions.
category: Interaction Classification
The category field groups related behaviors. Common values in evals/cases.jsonl include:
"direct-answer"— Simple information requests"debugging"— Technical troubleshooting scenarios"safety"— Content filtering and harm prevention tests
Categories enable targeted test runs: filter by category to validate only safety-critical paths.
prompt: Model Input
The prompt field contains the raw text presented to the evaluated model. This may include:
- Multi-turn conversation history
- System instructions
- Deliberately ambiguous or edge-case phrasing
case = {
"prompt": "What's 100 + 2? Give only the number, no explanation.",
...
}
risk: Safety Prioritization
The risk field drives evaluation urgency and human review workflows:
"low"— Standard functional tests"medium"— Potential policy edge cases"high"— Safety-critical scenarios requiring immediate attention
Filter high-priority cases using jq:
jq -c 'select(.risk=="high")' evals/cases.jsonl
criteria: Grading Rubric
The criteria array defines explicit pass conditions. Each string is a convertible assertion.
case = {
"criteria": [
"Answers 102.",
"Does not invent unnecessary steps for the user.",
"Avoids asking clarifying questions."
]
}
Evaluation scripts iterate this array and test model outputs against each criterion.
Practical Parsing Examples
Python Line-by-Line Processing
import json
def stream_cases(path='evals/cases.jsonl'):
"""Yield case objects without loading full file."""
with open(path, 'r') as f:
for line_num, line in enumerate(f, 1):
try:
yield json.loads(line)
except json.JSONDecodeError as e:
print(f"Invalid JSON at line {line_num}: {e}")
# Inspect high-risk debugging cases
for case in stream_cases():
if case['risk'] == 'high' and case['category'] == 'debugging':
print(f"{case['id']}: {case['prompt'][:60]}...")
for criterion in case['criteria']:
print(f" - {criterion}")
Node.js Streaming Parser
const fs = require('fs');
const readline = require('readline');
async function* caseGenerator(path = 'evals/cases.jsonl') {
const stream = fs.createReadStream(path);
const rl = readline.createInterface({ input: stream, crlfDelay: Infinity });
for await (const line of rl) {
if (!line.trim()) continue;
yield JSON.parse(line);
}
}
// Usage: count cases by risk level
async function riskSummary() {
const counts = { low: 0, medium: 0, high: 0 };
for await (const caseObj of caseGenerator()) {
counts[caseObj.risk] = (counts[caseObj.risk] || 0) + 1;
}
console.table(counts);
}
riskSummary();
Batch Validation with jq
Extract specific fields for CSV export:
jq -r '[.id, .risk, .category, (.criteria | length)] | @csv' evals/cases.jsonl
Find cases with empty criteria arrays (data quality check):
jq 'select((.criteria | length) == 0)' evals/cases.jsonl
Repository Integration Points
The evals/cases.jsonl file connects to several components in ayghri/i-have-adhd:
scripts/run_evals.py— Consumes this file to orchestrate evaluation runstests/directory — Unit tests that validate case structure and grading logic- CI pipelines — Load subsets of cases based on
riskfield for staged testing
When adding new evaluation scenarios, maintain the five-field schema and validate with:
# Structural validation: ensure all required fields exist
jq 'select(
(.id | type) != "string" or
(.category | type) != "string" or
(.prompt | type) != "string" or
(.risk | type) != "string" or
(.criteria | type) != "array"
) | "INVALID: " + .id' evals/cases.jsonl
Summary
evals/cases.jsonluses JSON Lines format—one independent JSON object per line for streaming efficiency- Each case contains five required fields:
id(string),category(string),prompt(string),risk(string), andcriteria(array of strings) - The
riskfield enables priority filtering:"low","medium", or"high" - The
criteriaarray defines pass/fail conditions for automated grading - Python (
jsonmodule), Node.js (readline), andjqall parse this format efficiently
Frequently Asked Questions
What is the difference between category and risk in evals/cases.jsonl?
The category field describes the type of interaction (e.g., debugging, direct-answer, safety), while risk indicates the severity level for safety review. A single category like "debugging" can contain low-risk formatting questions and high-risk code execution scenarios. Use category for test organization and risk for prioritization.
How do I add a new evaluation case to the repository?
Create a JSON object with all five required fields and append it as one line to evals/cases.jsonl. Ensure valid JSON with no trailing commas, and include at least one criterion in the criteria array. Run structural validation with jq before committing.
Why use JSON Lines instead of a standard JSON array?
JSON Lines enables streaming processing—evaluation scripts can process cases one at a time without loading the entire file into memory. This matters when running thousands of cases against API-rate-limited models. It also allows tail, head, and grep operations on the raw file.
How are failing criteria detected during evaluation?
The scripts/run_evals.py script (per the repository structure) loads each case and sends the prompt to the model under test. The response is evaluated against each string in the criteria array—typically through pattern matching, LLM-as-judge calls, or assertion functions. Results aggregate by id for debugging and reporting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →