# How to Use the Release Gate (`soup ship`) for Model Quality Validation

> Validate model quality for production using the soup ship release gate. Ensure task improvement and prevent catastrophic forgetting with this essential tool.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**The `soup ship` command is the built-in release gate that validates whether a fine-tuned model is ready for production by checking task improvement and guarding against catastrophic forgetting.**

The `soup ship` release gate in the **Soup** repository provides a deterministic, CI-friendly mechanism for model quality validation. This two-leg verdict engine ensures your fine-tuned model improves on its target task without regressing on general capabilities. All core logic lives in pure Python without GPU dependencies, making it fast to run in any pipeline.

## Understanding the Two-Leg Verdict Engine

The validation system implemented in [`src/soup_cli/utils/ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ship_verdict.py) evaluates models through two independent checks that must both pass.

### Leg 1: Task Win Validation

The **task win** verifies that your tuned model outperforms the base model on the specific task you optimized for. The function `build_task_win` constructs a `TaskWin` dataclass containing:

- `mode`: The evaluation protocol (`metric`, `judge_score`, or `pairwise`)
- `base`: Base model score
- `tuned`: Tuned model score
- `won`: Boolean indicating strict improvement (`tuned > base`)

The supported modes are defined in `TASK_MODES` and `SUPPORTED_TASK_MODES` within the same file. The comparison is **strict**—equal scores result in rejection.

### Leg 2: Catastrophic Forgetting Guard

The **general-suite benchmark** check prevents regression on previously learned capabilities. For each benchmark, `compute_benchmark_deltas` creates a `BenchmarkDelta` tracking:

- `base`: Original score
- `tuned`: New score
- `delta`: Calculated difference
- `regressed`: Flag when `base - tuned` exceeds the forgetting threshold

The default **forgetting threshold** is `0.05` (`DEFAULT_FORGETTING_THRESHOLD`). To avoid false positives from instrument noise, the system also measures a **noise floor** via `compute_noise_floor`, with run bounds set by `MIN_NOISE_FLOOR_RUNS` and `MAX_NOISE_FLOOR_RUNS`.

### Final Verdict Logic

The `decide_ship` function combines both legs:

```

SHIP   ⇔  (task_win.won == True)  AND  (no benchmark regressed)
DON'T SHIP otherwise

```

Failed rules receive specific codes for debugging: `FAILED_MISSING_BASELINE`, `FAILED_TASK_WIN`, or `FAILED_REGRESSION`.

## Command-Line Usage

The CLI wrapper in [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py) handles argument parsing, flag validation, and evidence management.

### Basic Validation Run

```bash
soup ship \
  --base base_model_dir \
  --adapter lora_adapter_dir \
  --task-eval tasks.jsonl \
  --task-mode metric

```

### Offline Mode with Pre-Computed Evidence

```bash
soup ship --evidence eval_results.json \
          --output verdict.json

```

### Emit Evidence for Reproducibility

```bash
soup ship --base base_dir \
          --adapter lora_dir \
          --task-eval tasks.jsonl \
          --emit-evidence verdict.ev.json \
          --push myorg/myrepo#42

```

### Key CLI Flags

| Flag | Purpose | Default |
|------|---------|---------|
| `--forgetting-threshold` | Maximum allowed regression | `0.05` |
| `--noise-floor` | Enable noise-floor calculation | auto-detected |
| `--task-mode` | Evaluation protocol | `metric` |
| `--evidence` | Load pre-computed results | none |
| `--emit-evidence` | Save serialized verdict | none |

## Exit Codes for CI Integration

The command returns structured exit codes for pipeline automation:

- **0** — SHIP (model passes validation)
- **2** — DON'T SHIP (regression or task-win failure)
- **3** — Usage/validation error
- **1** — Runtime error (I/O, model load failure)

## Programmatic API

Import the core engine for custom validation workflows:

```python
from soup_cli.utils.ship_verdict import (
    build_task_win,
    compute_benchmark_deltas,
    decide_ship,
    verdict_to_dict,
)

# Build task win evaluation

task_win = build_task_win(mode="metric", base=0.73, tuned=0.78)

# Compute benchmark deltas

deltas = compute_benchmark_deltas({
    "arc": (0.6, 0.55),       # (base, tuned)

    "hellaswag": (0.71, 0.71)
})

# Generate verdict

verdict = decide_ship(task_win, deltas)

print(verdict.decision)           # "SHIP" or "DON'T SHIP"

print(verdict_to_dict(verdict))   # serializable dictionary

```

## Output and Visualization

The `render_ship_panel` function produces a Rich-styled terminal display including:

- Decision status with color coding
- One-line summary
- Markdown-compatible evidence block for documentation

Control characters are sanitized via `for_terminal` before printing.

## Architectural Safety Features

The `soup ship` release gate includes multiple safeguards:

1. **Pure-Python core** — No top-level Torch imports; runs CPU-only without GPU requirements
2. **Input size caps** — `_MAX_EVIDENCE_BYTES` and `_MAX_SUITE_BENCHMARKS` prevent DoS attacks
3. **Replayable evidence** — `verdict_to_evidence` serializes decisions to JSON for deterministic re-runs via `--evidence`
4. **Noise floor measurement** — Statistical validation of base model variance stored in `NoiseFloor` dataclass

## Key Source Files

| File | Role |
|------|------|
| [`src/soup_cli/utils/ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ship_verdict.py) | Core verdict engine, dataclasses (`TaskWin`, `BenchmarkDelta`, `NoiseFloor`, `ShipVerdict`), noise-floor logic, and rendering |
| [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py) | CLI entry point, flag validation, evidence I/O, exit code handling |
| [`tests/test_eval_gate.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_eval_gate.py) | Integration tests for realistic benchmark suites |
| [`docs/commands.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/commands.md) | Human-readable command reference |

## Summary

- **Two-leg validation** ensures task improvement (`build_task_win`) and prevents catastrophic forgetting (`compute_benchmark_deltas`)
- **Strict comparison** requires `tuned > base`; equal scores fail the gate
- **Configurable thresholds** via `--forgetting-threshold` (default 0.05) and noise-floor detection
- **CI-native design** with structured exit codes (0/2/3/1) and JSON evidence serialization
- **GPU-free operation** enables fast, reproducible validation in any environment

## Frequently Asked Questions

### What happens if my tuned model matches the base model score exactly?

The task win uses **strict greater-than comparison** (`tuned > base`). Equal scores result in `FAILED_TASK_WIN` and a DON'T SHIP verdict. You must demonstrate measurable improvement to pass the release gate.

### Can I run validation without loading models into memory?

Yes. Use the **offline mode** with `--evidence eval_results.json` to validate pre-computed scores. This bypasses model loading entirely and operates on JSON data, making it suitable for shared CI runners without GPU access.

### How does the noise floor prevent false regression alerts?

The noise floor (`compute_noise_floor`) measures base model variance through repeated runs. A benchmark delta only counts as regression if it exceeds **both** the forgetting threshold AND the measured noise floor. This prevents rejecting models due to evaluation instrument instability rather than actual capability degradation.

### What evidence format does `--emit-evidence` produce?

The `verdict_to_evidence` function outputs a JSON file containing the complete `ShipVerdict` dataclass: task win details, all benchmark deltas, noise floor measurements, failed rule codes, and the final decision. This file can be fed back via `--evidence` for identical re-runs or attached to pull requests for audit trails.