How to Use the Release Gate (`soup ship`) for Model Quality Validation
The soup ship command is the built-in release gate that validates whether a fine-tuned model is ready for production by checking task improvement and guarding against catastrophic forgetting.
The soup ship release gate in the Soup repository provides a deterministic, CI-friendly mechanism for model quality validation. This two-leg verdict engine ensures your fine-tuned model improves on its target task without regressing on general capabilities. All core logic lives in pure Python without GPU dependencies, making it fast to run in any pipeline.
Understanding the Two-Leg Verdict Engine
The validation system implemented in src/soup_cli/utils/ship_verdict.py evaluates models through two independent checks that must both pass.
Leg 1: Task Win Validation
The task win verifies that your tuned model outperforms the base model on the specific task you optimized for. The function build_task_win constructs a TaskWin dataclass containing:
mode: The evaluation protocol (metric,judge_score, orpairwise)base: Base model scoretuned: Tuned model scorewon: Boolean indicating strict improvement (tuned > base)
The supported modes are defined in TASK_MODES and SUPPORTED_TASK_MODES within the same file. The comparison is strict—equal scores result in rejection.
Leg 2: Catastrophic Forgetting Guard
The general-suite benchmark check prevents regression on previously learned capabilities. For each benchmark, compute_benchmark_deltas creates a BenchmarkDelta tracking:
base: Original scoretuned: New scoredelta: Calculated differenceregressed: Flag whenbase - tunedexceeds the forgetting threshold
The default forgetting threshold is 0.05 (DEFAULT_FORGETTING_THRESHOLD). To avoid false positives from instrument noise, the system also measures a noise floor via compute_noise_floor, with run bounds set by MIN_NOISE_FLOOR_RUNS and MAX_NOISE_FLOOR_RUNS.
Final Verdict Logic
The decide_ship function combines both legs:
SHIP ⇔ (task_win.won == True) AND (no benchmark regressed)
DON'T SHIP otherwise
Failed rules receive specific codes for debugging: FAILED_MISSING_BASELINE, FAILED_TASK_WIN, or FAILED_REGRESSION.
Command-Line Usage
The CLI wrapper in src/soup_cli/commands/ship.py handles argument parsing, flag validation, and evidence management.
Basic Validation Run
soup ship \
--base base_model_dir \
--adapter lora_adapter_dir \
--task-eval tasks.jsonl \
--task-mode metric
Offline Mode with Pre-Computed Evidence
soup ship --evidence eval_results.json \
--output verdict.json
Emit Evidence for Reproducibility
soup ship --base base_dir \
--adapter lora_dir \
--task-eval tasks.jsonl \
--emit-evidence verdict.ev.json \
--push myorg/myrepo#42
Key CLI Flags
| Flag | Purpose | Default |
|---|---|---|
--forgetting-threshold |
Maximum allowed regression | 0.05 |
--noise-floor |
Enable noise-floor calculation | auto-detected |
--task-mode |
Evaluation protocol | metric |
--evidence |
Load pre-computed results | none |
--emit-evidence |
Save serialized verdict | none |
Exit Codes for CI Integration
The command returns structured exit codes for pipeline automation:
- 0 — SHIP (model passes validation)
- 2 — DON'T SHIP (regression or task-win failure)
- 3 — Usage/validation error
- 1 — Runtime error (I/O, model load failure)
Programmatic API
Import the core engine for custom validation workflows:
from soup_cli.utils.ship_verdict import (
build_task_win,
compute_benchmark_deltas,
decide_ship,
verdict_to_dict,
)
# Build task win evaluation
task_win = build_task_win(mode="metric", base=0.73, tuned=0.78)
# Compute benchmark deltas
deltas = compute_benchmark_deltas({
"arc": (0.6, 0.55), # (base, tuned)
"hellaswag": (0.71, 0.71)
})
# Generate verdict
verdict = decide_ship(task_win, deltas)
print(verdict.decision) # "SHIP" or "DON'T SHIP"
print(verdict_to_dict(verdict)) # serializable dictionary
Output and Visualization
The render_ship_panel function produces a Rich-styled terminal display including:
- Decision status with color coding
- One-line summary
- Markdown-compatible evidence block for documentation
Control characters are sanitized via for_terminal before printing.
Architectural Safety Features
The soup ship release gate includes multiple safeguards:
- Pure-Python core — No top-level Torch imports; runs CPU-only without GPU requirements
- Input size caps —
_MAX_EVIDENCE_BYTESand_MAX_SUITE_BENCHMARKSprevent DoS attacks - Replayable evidence —
verdict_to_evidenceserializes decisions to JSON for deterministic re-runs via--evidence - Noise floor measurement — Statistical validation of base model variance stored in
NoiseFloordataclass
Key Source Files
| File | Role |
|---|---|
src/soup_cli/utils/ship_verdict.py |
Core verdict engine, dataclasses (TaskWin, BenchmarkDelta, NoiseFloor, ShipVerdict), noise-floor logic, and rendering |
src/soup_cli/commands/ship.py |
CLI entry point, flag validation, evidence I/O, exit code handling |
tests/test_eval_gate.py |
Integration tests for realistic benchmark suites |
docs/commands.md |
Human-readable command reference |
Summary
- Two-leg validation ensures task improvement (
build_task_win) and prevents catastrophic forgetting (compute_benchmark_deltas) - Strict comparison requires
tuned > base; equal scores fail the gate - Configurable thresholds via
--forgetting-threshold(default 0.05) and noise-floor detection - CI-native design with structured exit codes (0/2/3/1) and JSON evidence serialization
- GPU-free operation enables fast, reproducible validation in any environment
Frequently Asked Questions
What happens if my tuned model matches the base model score exactly?
The task win uses strict greater-than comparison (tuned > base). Equal scores result in FAILED_TASK_WIN and a DON'T SHIP verdict. You must demonstrate measurable improvement to pass the release gate.
Can I run validation without loading models into memory?
Yes. Use the offline mode with --evidence eval_results.json to validate pre-computed scores. This bypasses model loading entirely and operates on JSON data, making it suitable for shared CI runners without GPU access.
How does the noise floor prevent false regression alerts?
The noise floor (compute_noise_floor) measures base model variance through repeated runs. A benchmark delta only counts as regression if it exceeds both the forgetting threshold AND the measured noise floor. This prevents rejecting models due to evaluation instrument instability rather than actual capability degradation.
What evidence format does --emit-evidence produce?
The verdict_to_evidence function outputs a JSON file containing the complete ShipVerdict dataclass: task win details, all benchmark deltas, noise floor measurements, failed rule codes, and the final decision. This file can be fed back via --evidence for identical re-runs or attached to pull requests for audit trails.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →