How the Scorecard Outcome Grading System Works: Programmatic and LLM-as-Judge Evaluation
The Scorecard outcome grading system evaluates agent task completion using hybrid graders—deterministic Python functions for exact validation and LLM-as-judge for nuanced assessment—returning standardized (status, why) tuples that populate the final Scorecard JSON.
The Scorecard outcome grading system with programmatic and LLM-as-judge evaluation powers the agent evaluation pipeline in the anthropics/cwc-workshops repository. This architecture validates AI agent outputs against task definitions using either deterministic code checks or Claude-powered judgment, producing consistent pass/fail outcomes for the research-desk UI.
Core Architecture
The evaluation pipeline centers on the grade(task, result) function located in agent-decomposition/evals/graders.py. When an agent finishes a task, it produces a result object containing the final generated text, actions taken, token usage, and turn count. A task definition specifies which grader to apply and what ground-truth values to compare against.
Every grader in the system conforms to the same contract: returning a tuple (status, why) where status is pass, fail, or pass-slow, and why provides an explanation for non-passing states.
Programmatic Graders
Programmatic graders are pure Python functions that inspect result objects without external API calls. These graders perform deterministic checks defined in agent-decomposition/evals/graders.py.
Available programmatic graders include:
exact_match– Validates exact string or numeric equality against expected valuesset_match– Compares unordered collections (e.g., required SKU sets)numeric_tolerance– Checks if values fall within a specified percentage rangeaction_taken– Verifies specific tool calls or actions appear in the resultefficiency– Validates token usage and turn counts stay within budgetswall_budget– Enforces time-based execution limitsranked_mention– Checks for ordered appearance of key terms
These functions parse the result object directly and return immediate boolean-style outcomes without LLM latency.
LLM-as-Judge Evaluation
The LLM-as-judge grader leverages anthropic.Anthropic() to apply nuanced, rubric-based evaluation when deterministic rules are insufficient. Implemented in the llm_judge function within agent-decomposition/evals/graders.py, this grader sends the agent's final text plus a task-defined rubric to Claude.
The model analyzes the content against criteria such as tone, completeness, or reasoning quality, then responds with "PASS: ..." or "FAIL: ..." prefixes. The wrapper parses this response and maps it to the standard (status, why) format, allowing natural language criteria to integrate seamlessly with programmatic checks.
Composite Grading
The composite grader enables AND-combination of multiple sub-graders. Defined in agent-decomposition/evals/graders.py, this meta-grader accepts a list of checks and applies them sequentially.
If any sub-grader returns fail, the composite returns fail. If all pass but any return pass-slow, the overall status downgrades to pass-slow. Only when all sub-graders return pass does the composite return pass. This allows complex validation rules that mix precise numeric thresholds with subjective quality assessment.
Implementation Examples
Basic Programmatic Check
Validate exact numeric output and SKU matching:
from agent_decomposition.evals.graders import grade
task = {
"grader": "exact_match",
"expected": {"source": "reorder_qty", "sku": "SKU-1234", "qty": 150},
"budget_turns": 10,
"budget_tokens": 5000,
}
status, why = grade(task, result)
print(status) # → "pass" or "fail"
print(why) # → "" or failure reason
Implementation reference: graders.py lines 47-51.
LLM Rubric Evaluation
Apply subjective quality criteria using Claude:
task = {
"grader": "llm_judge",
"expected": {
"rubric": """
1. The response must mention the correct SKU.
2. Quantity must be within ±20% of the forecast.
3. The tone should be professional.
"""
}
}
status, why = grade(task, result)
Implementation reference: llm_judge function in graders.py lines 16-30.
Combining Multiple Graders
Enforce both numeric accuracy and stylistic requirements:
task = {
"grader": "composite",
"expected": {
"checks": [
{"grader": "numeric_tolerance", "tolerance_pct": 20},
{"grader": "llm_judge", "rubric": "Must be concise and professional."}
]
}
}
status, why = grade(task, result)
Implementation reference: composite function in graders.py lines 2-14.
Key Source Files
The grading pipeline spans Python evaluation logic and TypeScript schema validation:
agent-decomposition/evals/graders.py– Core grading registry (GRADERSmap) and all grader implementations (exact_match,llm_judge,composite, etc.)agent-decomposition/evals/report.py– Aggregates graded outcomes and writes the final Scorecard JSON tooutputs/<TICKER>/scorecard.jsonresearch-desk/src/lib/scorecard.ts– JSON schema definition andvalidateScorecardfunction ensuring outputs meet ticker, guidance tone, and confidence field requirementsresearch-desk/src/lib/analysis.ts– Orchestrates agent runs, invokesgrade(), and aggregates results into the Scorecard structure consumed by the UIresearch-desk/src/components/ScorecardsPanel.tsx– React component displaying validated Scorecards in the research-desk interface
Summary
The Scorecard outcome grading system combines deterministic validation with flexible AI judgment:
- Standardized interface – All graders return
(status, why)tuples enabling modular mixing of check types - Dual evaluation modes – Programmatic graders for precision, LLM-as-judge for nuance
- Composable logic –
compositegrader chains multiple validators with AND semantics andpass-slowpropagation - Schema enforcement – Pre-grading validation in
scorecard.tsensures structural compliance before evaluation - UI integration – Results flow through
report.pyinto persisted JSON consumed by the Scorecards panel
Frequently Asked Questions
How does the LLM-as-judge grader handle ambiguous rubric criteria?
The llm_judge function passes the rubric directly to Claude with explicit instructions to prefix responses with "PASS:" or "FAIL:". The wrapper parses these prefixes deterministically, converting subjective model assessments into the binary pass/fail status required by the pipeline while preserving the model's explanation in the why field.
Can custom graders be added to the existing registry?
Yes. New graders must accept (task, result) parameters and return (status, why) tuples. Register them in the GRADERS dictionary in agent-decomposition/evals/graders.py, mapping a string key to the function. Once registered, reference the key in task definitions via the "grader" field.
What triggers a pass-slow status outcome?
The pass-slow status indicates functional correctness with performance penalties. Programmatic graders like efficiency or wall_budget return this when agents exceed token counts, turn limits, or time budgets. In composite grading, if any sub-grader returns pass-slow while others pass, the aggregate result becomes pass-slow.
Where does the final Scorecard data persist?
After grading completes, agent-decomposition/evals/report.py writes outcomes to outputs/<TICKER>/scorecard.json. The research-desk/src/lib/analysis.ts module reads these files to populate the Scorecards tab, while scorecard.ts validates schema compliance before UI rendering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →