How the `--noise-floor` Option in `soup ship` Handles GPU Non-Determinism

The --noise-floor flag quantifies run-to-run variability from GPU non-determinism by re-executing the base model multiple times and aggregating results into a statistical noise floor.

Running inference on GPUs—even with greedy decoding—produces non-deterministic outputs due to hardware-level variations. The Soup CLI's ship command provides the --noise-floor option to measure this intrinsic variance, giving users a baseline for evaluating actual model improvements versus measurement noise.

What GPU Non-Determinism Looks Like in Practice

The Soup codebase explicitly acknowledges this behavior. As noted in src/soup_cli/utils/ship_verdict.py at lines 102-105, "Greedy decoding is not deterministic on GPU. Measured on an H100, the same..." This documentation confirms that identical inputs, models, and decoding strategies can yield different scores across repeated runs on identical hardware.

Without quantifying this noise, users might misinterpret natural score fluctuations as meaningful performance changes when comparing models or configurations.

How --noise-floor Works: Implementation Details

Command Parsing and Validation

The flag enters the system through src/soup_cli/commands/ship.py. The function _validate_noise_floor_flag parses user input and ensures the repeat count is valid before passing it as noise_floor_runs to _verdict_live:


# In src/soup_cli/commands/ship.py

noise_floor = _validate_noise_floor_flag(noise_floor)          # parses the flag

...
verdict = _verdict_live(
    ...,
    noise_floor_runs=noise_floor,  # passes repeat count to the core logic

)

Core Purpose: Isolating GPU-Induced Variance

The implementation specifically targets what lines 746-755 describe: "noise-floor repeat so the spread reflects run-to-run variance, not setup." This distinction matters because:

  • Without noise floor measurement: Score differences might stem from configuration changes, model changes, or GPU randomness
  • With noise floor measurement: Users see a statistical spread representing purely GPU-induced variation, enabling cleaner comparisons

Execution Flow

When --noise-floor N is specified:

  1. The base model runs N times on bundled evaluation suites
  2. Each run performs full inference unchanged from the base configuration
  3. Scores are collected and aggregated into variance statistics
  4. Results integrate into the final verdict as a measurable "noise floor"

Critical Limitations and Constraints

Bundled Suites Only

The flag carries an important scope restriction. As implemented in src/soup_cli/commands/ship.py at lines 829-831, --noise-floor only works with bundled suites. Attempting to use it with custom suites triggers a warning and the flag is ignored.

This limitation exists because bundled suites have standardized evaluation patterns that support repeatable measurement, while custom suites may lack deterministic structure.

Practical Usage Examples

Basic Evaluation Without Noise Floor


# Single base model evaluation

soup ship --base my-base-model --task-eval my-task

Measuring GPU Non-Determinism


# Capture noise floor with 5 repeats to quantify H100-level variance

soup ship --base my-base-model --task-eval my-task --noise-floor 5

Higher repeat counts yield more stable noise floor estimates at the cost of increased compute time.

Result Integration and Interpretation

The computed noise-floor statistics flow into the final verdict display. Users receive:

  • Absolute scores from single or repeated runs
  • Variance/spread metrics representing GPU-induced fluctuation range
  • Context for improvement assessment—changes smaller than the noise floor lack statistical significance

This integration happens in src/soup_cli/utils/ship_verdict.py, which implements the repeat execution loop and statistical aggregation.

Key Source Files

File Role
src/soup_cli/commands/ship.py CLI entry point; parses --noise-floor, validates input, forwards repeat count
src/soup_cli/utils/ship_verdict.py Executes repeated inference, collects scores, computes noise-floor statistics
src/soup_cli/config/schema.py Documents CLI validator, notes GPU non-determinism relationship

Summary

  • --noise-floor N triggers N repeated base model runs to measure GPU-induced variance
  • The flag quantifies non-determinism explicitly acknowledged for H100 and similar GPUs during greedy decoding
  • Results provide a statistical baseline for distinguishing real improvements from measurement noise
  • Bundled suites only—custom suites trigger warnings and skip noise-floor calculation
  • Implementation spans ship.py for CLI handling and ship_verdict.py for statistical computation

Frequently Asked Questions

Why is greedy decoding non-deterministic on GPU?

GPU hardware executes operations in parallel with floating-point rounding variations, memory access timing differences, and kernel scheduling non-determinism. Even without random sampling, these factors create measurable output variance across runs—documented in the Soup codebase specifically for H100 hardware.

What happens if I set --noise-floor with a custom evaluation suite?

The system emits a warning and ignores the flag. As enforced in src/soup_cli/commands/ship.py lines 829-831, noise-floor measurement requires bundled suites with standardized structure. Custom suites lack guaranteed determinism in their evaluation patterns.

How many repeats should I specify for accurate noise floor measurement?

The codebase doesn't prescribe a default, but statistical reliability typically requires 5-10 repeats minimum for variance estimation, with 20+ repeats for precise confidence intervals. Balance accuracy needs against compute costs—each repeat executes full base model inference.

Does --noise-floor affect the ship verdict comparison between base and shipped models?

Yes—indirectly. The noise floor doesn't change model scores but provides interpretation context. When comparing base versus shipped model performance, improvements smaller than the measured noise floor likely represent measurement artifact rather than genuine capability differences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →