How to Set Up Eval-Gated Training with the SHIP Verdict in Soup
Eval-gated training automatically halts fine-tuning when the SHIP verdict engine detects a benchmark regression, using the same regression threshold logic as the soup ship command.
The Soup framework provides a built-in safety mechanism that runs declarative evaluation suites at epoch boundaries and stops training when performance drops exceed your configured threshold. This guide covers configuration via YAML, CLI shortcuts, and programmatic Python APIs based on the actual implementation in MakazhanAlpamys/Soup.
What Eval-Gated Training Does
Eval-gated training integrates the SHIP verdict directly into the training loop. After each epoch (or custom interval), Soup evaluates your model against a benchmark suite, compares results to a baseline, and applies a policy action if any metric regresses beyond the threshold.
The gate uses identical regression semantics to soup ship: a benchmark is regressed when base_score - tuned_score > regression_threshold. The default threshold is 0.05 (5%).
Core Configuration: EvalGateConfig
All gate parameters are defined by the EvalGateConfig Pydantic model in src/soup_cli/config/schema.py (lines 668-704):
| Parameter | Type | Description |
|---|---|---|
enabled |
bool |
Master switch for the gate |
suite |
str |
Path to YAML evaluation suite |
every_n_epochs |
int |
Evaluation frequency (default: 1) |
regression_threshold |
float |
Maximum allowed score drop (default: 0.05) |
baseline |
str |
"previous" | "original" | path to checkpoint |
on_regression |
str |
"stop" (default) | "warn" | "continue" |
Method 1: Configure via YAML Config File
Create a soup_config.yaml with the training.eval_gate block:
training:
epochs: 10
eval_gate:
enabled: true
suite: evals/gate.yaml
every_n_epochs: 1
regression_threshold: 0.05
baseline: previous
on_regression: stop
Run training:
soup train --config soup_config.yaml
The suite path resolves against Soup's bundled suites in src/soup_cli/eval/gate_suites.py or your custom filesystem paths.
Method 2: Use the CLI --gate Shortcut
For quick activation, the train command provides a compact flag expansion. The --gate <suite> shortcut auto-populates the full eval_gate configuration:
soup train --gate evals/gate.yaml \
--training.epochs 10 \
--training.eval_gate.enabled true
The expansion logic lives in src/soup_cli/commands/train.py (lines 681-692), where the parser injects:
enabled: truesuite: <provided_path>every_n_epochs: 1(default)regression_threshold: 0.05(default)baseline: previous(default)on_regression: stop(default)
Method 3: Programmatic Python Configuration
Instantiate EvalGateConfig directly for dynamic training scripts:
from soup_cli.config.schema import EvalGateConfig, TrainingConfig
# Configure the gate
gate_cfg = EvalGateConfig(
enabled=True,
suite="evals/gate.yaml",
every_n_epochs=2,
regression_threshold=0.04,
baseline="previous",
on_regression="warn", # Log warning but continue training
)
# Attach to training configuration
train_cfg = TrainingConfig(
epochs=20,
eval_gate=gate_cfg,
# ... other options ...
)
# Any trainer using SoupTrainerCallback will respect the gate
from soup_cli.trainer.sft import SoupTrainer
trainer = SoupTrainer(config=train_cfg, model=model, datasets=datasets)
trainer.train() # Automatically halts/warns per policy
How the Trainer Executes the Gate
The SoupTrainerCallback in src/soup_cli/monitoring/callback.py implements the gating logic:
-
Initialization (lines 692-733): On first epoch, the callback:
- Loads the evaluation suite from YAML
- Resolves the baseline checkpoint source
- Builds a lightweight generation function for evaluation
-
Epoch execution (lines 71-84):
._run_eval_gate()runs the suite and applies the policy:- Scores each benchmark in the suite
- Computes regression against baseline
- If any benchmark exceeds
regression_threshold, takeson_regressionaction
The same ship_verdict.py utilities power both the training gate and the standalone soup ship command, ensuring consistent regression detection.
Regression Policies Explained
| Policy | Behavior | Use Case |
|---|---|---|
stop |
Immediately terminate training (default) | Production fine-tuning where regressions are unacceptable |
warn |
Log to console/wandb, continue training | Experimental runs where you want visibility without interruption |
continue |
Silent continuation | Debugging or when gate is temporarily disabled |
Summary
- Eval-gated training integrates SHIP verdict regression detection directly into the training loop via
EvalGateConfig - Configure through YAML, the
--gateCLI shortcut, or Python instantiation - The gate executes in
SoupTrainerCallbackat epoch boundaries, comparing live results againstbaselineusingregression_threshold - All regression logic is shared with
soup shipinsrc/soup_cli/utils/ship_verdict.py
Frequently Asked Questions
What evaluation suites can I use with the eval gate?
Any valid Soup evaluation suite YAML works. Soup bundles predefined suites in src/soup_cli/eval/gate_suites.py—reference these by name like evals/gate.yaml or provide absolute paths to custom suites. The suite format matches standard Soup evaluation specifications.
How does the eval gate differ from running soup ship after training?
The eval gate automates the same regression check that soup ship performs manually. Both use identical threshold logic from ship_verdict.py, but the gate runs automatically at epoch boundaries and can stop training immediately—preventing wasted compute on regressing models rather than discovering problems post-hoc.
Can I change the gate frequency to run every N steps instead of epochs?
The every_n_epochs parameter controls execution frequency. Soup does not currently support step-level gating; the gate triggers at epoch boundaries to ensure complete model state evaluation. For finer granularity, reduce batch size to increase epoch frequency or use fractional epoch pseudo-scheduling.
What happens if the baseline checkpoint is missing?
The baseline parameter accepts "previous" (last saved checkpoint), "original" (pre-training weights), or explicit paths. If a baseline cannot be resolved, the SoupTrainerCallback initialization (lines 692-733) raises a clear configuration error before training begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →