How to Use the Release Gate (`soup ship`) for Model Quality Validation

The soup ship command is the built-in release gate that validates whether a fine-tuned model is ready for production by checking task improvement and guarding against catastrophic forgetting.

The soup ship release gate in the Soup repository provides a deterministic, CI-friendly mechanism for model quality validation. This two-leg verdict engine ensures your fine-tuned model improves on its target task without regressing on general capabilities. All core logic lives in pure Python without GPU dependencies, making it fast to run in any pipeline.

Understanding the Two-Leg Verdict Engine

The validation system implemented in src/soup_cli/utils/ship_verdict.py evaluates models through two independent checks that must both pass.

Leg 1: Task Win Validation

The task win verifies that your tuned model outperforms the base model on the specific task you optimized for. The function build_task_win constructs a TaskWin dataclass containing:

  • mode: The evaluation protocol (metric, judge_score, or pairwise)
  • base: Base model score
  • tuned: Tuned model score
  • won: Boolean indicating strict improvement (tuned > base)

The supported modes are defined in TASK_MODES and SUPPORTED_TASK_MODES within the same file. The comparison is strict—equal scores result in rejection.

Leg 2: Catastrophic Forgetting Guard

The general-suite benchmark check prevents regression on previously learned capabilities. For each benchmark, compute_benchmark_deltas creates a BenchmarkDelta tracking:

  • base: Original score
  • tuned: New score
  • delta: Calculated difference
  • regressed: Flag when base - tuned exceeds the forgetting threshold

The default forgetting threshold is 0.05 (DEFAULT_FORGETTING_THRESHOLD). To avoid false positives from instrument noise, the system also measures a noise floor via compute_noise_floor, with run bounds set by MIN_NOISE_FLOOR_RUNS and MAX_NOISE_FLOOR_RUNS.

Final Verdict Logic

The decide_ship function combines both legs:


SHIP   ⇔  (task_win.won == True)  AND  (no benchmark regressed)
DON'T SHIP otherwise

Failed rules receive specific codes for debugging: FAILED_MISSING_BASELINE, FAILED_TASK_WIN, or FAILED_REGRESSION.

Command-Line Usage

The CLI wrapper in src/soup_cli/commands/ship.py handles argument parsing, flag validation, and evidence management.

Basic Validation Run

soup ship \
  --base base_model_dir \
  --adapter lora_adapter_dir \
  --task-eval tasks.jsonl \
  --task-mode metric

Offline Mode with Pre-Computed Evidence

soup ship --evidence eval_results.json \
          --output verdict.json

Emit Evidence for Reproducibility

soup ship --base base_dir \
          --adapter lora_dir \
          --task-eval tasks.jsonl \
          --emit-evidence verdict.ev.json \
          --push myorg/myrepo#42

Key CLI Flags

Flag Purpose Default
--forgetting-threshold Maximum allowed regression 0.05
--noise-floor Enable noise-floor calculation auto-detected
--task-mode Evaluation protocol metric
--evidence Load pre-computed results none
--emit-evidence Save serialized verdict none

Exit Codes for CI Integration

The command returns structured exit codes for pipeline automation:

  • 0 — SHIP (model passes validation)
  • 2 — DON'T SHIP (regression or task-win failure)
  • 3 — Usage/validation error
  • 1 — Runtime error (I/O, model load failure)

Programmatic API

Import the core engine for custom validation workflows:

from soup_cli.utils.ship_verdict import (
    build_task_win,
    compute_benchmark_deltas,
    decide_ship,
    verdict_to_dict,
)

# Build task win evaluation

task_win = build_task_win(mode="metric", base=0.73, tuned=0.78)

# Compute benchmark deltas

deltas = compute_benchmark_deltas({
    "arc": (0.6, 0.55),       # (base, tuned)

    "hellaswag": (0.71, 0.71)
})

# Generate verdict

verdict = decide_ship(task_win, deltas)

print(verdict.decision)           # "SHIP" or "DON'T SHIP"

print(verdict_to_dict(verdict))   # serializable dictionary

Output and Visualization

The render_ship_panel function produces a Rich-styled terminal display including:

  • Decision status with color coding
  • One-line summary
  • Markdown-compatible evidence block for documentation

Control characters are sanitized via for_terminal before printing.

Architectural Safety Features

The soup ship release gate includes multiple safeguards:

  1. Pure-Python core — No top-level Torch imports; runs CPU-only without GPU requirements
  2. Input size caps — _MAX_EVIDENCE_BYTES and _MAX_SUITE_BENCHMARKS prevent DoS attacks
  3. Replayable evidence — verdict_to_evidence serializes decisions to JSON for deterministic re-runs via --evidence
  4. Noise floor measurement — Statistical validation of base model variance stored in NoiseFloor dataclass

Key Source Files

File Role
src/soup_cli/utils/ship_verdict.py Core verdict engine, dataclasses (TaskWin, BenchmarkDelta, NoiseFloor, ShipVerdict), noise-floor logic, and rendering
src/soup_cli/commands/ship.py CLI entry point, flag validation, evidence I/O, exit code handling
tests/test_eval_gate.py Integration tests for realistic benchmark suites
docs/commands.md Human-readable command reference

Summary

  • Two-leg validation ensures task improvement (build_task_win) and prevents catastrophic forgetting (compute_benchmark_deltas)
  • Strict comparison requires tuned > base; equal scores fail the gate
  • Configurable thresholds via --forgetting-threshold (default 0.05) and noise-floor detection
  • CI-native design with structured exit codes (0/2/3/1) and JSON evidence serialization
  • GPU-free operation enables fast, reproducible validation in any environment

Frequently Asked Questions

What happens if my tuned model matches the base model score exactly?

The task win uses strict greater-than comparison (tuned > base). Equal scores result in FAILED_TASK_WIN and a DON'T SHIP verdict. You must demonstrate measurable improvement to pass the release gate.

Can I run validation without loading models into memory?

Yes. Use the offline mode with --evidence eval_results.json to validate pre-computed scores. This bypasses model loading entirely and operates on JSON data, making it suitable for shared CI runners without GPU access.

How does the noise floor prevent false regression alerts?

The noise floor (compute_noise_floor) measures base model variance through repeated runs. A benchmark delta only counts as regression if it exceeds both the forgetting threshold AND the measured noise floor. This prevents rejecting models due to evaluation instrument instability rather than actual capability degradation.

What evidence format does --emit-evidence produce?

The verdict_to_evidence function outputs a JSON file containing the complete ShipVerdict dataclass: task win details, all benchmark deltas, noise floor measurements, failed rule codes, and the final decision. This file can be fed back via --evidence for identical re-runs or attached to pull requests for audit trails.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →