# How the `--noise-floor` Option in `soup ship` Handles GPU Non-Determinism

> Learn how soup ships --noise-floor option addresses GPU non-determinism by re-executing models and aggregating results into a statistical noise floor for accurate variability quantification.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: internals
- Published: 2026-08-16

---

**The `--noise-floor` flag quantifies run-to-run variability from GPU non-determinism by re-executing the base model multiple times and aggregating results into a statistical noise floor.**

Running inference on GPUs—even with **greedy decoding**—produces non-deterministic outputs due to hardware-level variations. The Soup CLI's `ship` command provides the `--noise-floor` option to measure this intrinsic variance, giving users a baseline for evaluating actual model improvements versus measurement noise.

## What GPU Non-Determinism Looks Like in Practice

The Soup codebase explicitly acknowledges this behavior. As noted in [`src/soup_cli/utils/ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ship_verdict.py) at lines 102-105, "Greedy decoding is not deterministic on GPU. Measured on an H100, the same..." This documentation confirms that identical inputs, models, and decoding strategies can yield different scores across repeated runs on identical hardware.

Without quantifying this noise, users might misinterpret natural score fluctuations as meaningful performance changes when comparing models or configurations.

## How `--noise-floor` Works: Implementation Details

### Command Parsing and Validation

The flag enters the system through [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py). The function `_validate_noise_floor_flag` parses user input and ensures the repeat count is valid before passing it as `noise_floor_runs` to `_verdict_live`:

```python

# In src/soup_cli/commands/ship.py

noise_floor = _validate_noise_floor_flag(noise_floor)          # parses the flag

...
verdict = _verdict_live(
    ...,
    noise_floor_runs=noise_floor,  # passes repeat count to the core logic

)

```

### Core Purpose: Isolating GPU-Induced Variance

The implementation specifically targets what lines 746-755 describe: "noise-floor repeat so the spread reflects run-to-run variance, not setup." This distinction matters because:

- **Without noise floor measurement**: Score differences might stem from configuration changes, model changes, *or* GPU randomness
- **With noise floor measurement**: Users see a statistical spread representing purely GPU-induced variation, enabling cleaner comparisons

### Execution Flow

When `--noise-floor N` is specified:

1. The base model runs **N times** on bundled evaluation suites
2. Each run performs **full inference** unchanged from the base configuration
3. Scores are collected and aggregated into variance statistics
4. Results integrate into the final verdict as a measurable "noise floor"

## Critical Limitations and Constraints

### Bundled Suites Only

The flag carries an important scope restriction. As implemented in [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py) at lines 829-831, `--noise-floor` **only works with bundled suites**. Attempting to use it with custom suites triggers a warning and the flag is ignored.

This limitation exists because bundled suites have standardized evaluation patterns that support repeatable measurement, while custom suites may lack deterministic structure.

## Practical Usage Examples

### Basic Evaluation Without Noise Floor

```bash

# Single base model evaluation

soup ship --base my-base-model --task-eval my-task

```

### Measuring GPU Non-Determinism

```bash

# Capture noise floor with 5 repeats to quantify H100-level variance

soup ship --base my-base-model --task-eval my-task --noise-floor 5

```

Higher repeat counts yield more stable noise floor estimates at the cost of increased compute time.

## Result Integration and Interpretation

The computed noise-floor statistics flow into the final verdict display. Users receive:

- **Absolute scores** from single or repeated runs
- **Variance/spread metrics** representing GPU-induced fluctuation range
- **Context for improvement assessment**—changes smaller than the noise floor lack statistical significance

This integration happens in [`src/soup_cli/utils/ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ship_verdict.py), which implements the repeat execution loop and statistical aggregation.

## Key Source Files

| File | Role |
|------|------|
| [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py) | CLI entry point; parses `--noise-floor`, validates input, forwards repeat count |
| [`src/soup_cli/utils/ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ship_verdict.py) | Executes repeated inference, collects scores, computes noise-floor statistics |
| [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) | Documents CLI validator, notes GPU non-determinism relationship |

## Summary

- **`--noise-floor N`** triggers **N repeated base model runs** to measure GPU-induced variance
- The flag **quantifies non-determinism** explicitly acknowledged for H100 and similar GPUs during greedy decoding
- Results provide a **statistical baseline** for distinguishing real improvements from measurement noise
- **Bundled suites only**—custom suites trigger warnings and skip noise-floor calculation
- Implementation spans [`ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/ship.py) for CLI handling and [`ship_verdict.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/ship_verdict.py) for statistical computation

## Frequently Asked Questions

### Why is greedy decoding non-deterministic on GPU?

GPU hardware executes operations in parallel with floating-point rounding variations, memory access timing differences, and kernel scheduling non-determinism. Even without random sampling, these factors create measurable output variance across runs—documented in the Soup codebase specifically for H100 hardware.

### What happens if I set `--noise-floor` with a custom evaluation suite?

The system emits a warning and ignores the flag. As enforced in [`src/soup_cli/commands/ship.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ship.py) lines 829-831, noise-floor measurement requires bundled suites with standardized structure. Custom suites lack guaranteed determinism in their evaluation patterns.

### How many repeats should I specify for accurate noise floor measurement?

The codebase doesn't prescribe a default, but statistical reliability typically requires **5-10 repeats minimum** for variance estimation, with **20+ repeats** for precise confidence intervals. Balance accuracy needs against compute costs—each repeat executes full base model inference.

### Does `--noise-floor` affect the ship verdict comparison between base and shipped models?

Yes—indirectly. The noise floor doesn't change model scores but provides **interpretation context**. When comparing base versus shipped model performance, improvements smaller than the measured noise floor likely represent measurement artifact rather than genuine capability differences.