# How to Use A/B Testing to Compare Model Variants with the Soup Framework

> Learn to use A/B testing for model variants with the Soup framework. Implement mSPRT for early stopping and guaranteed error control to optimize your models.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**The Soup framework implements sequential A/B testing for model variants using the Mixture Sequential Probability Ratio Test (mSPRT), enabling early stopping with guaranteed Type-I and Type-II error control.**

A/B testing model variants is essential for production ML systems, but traditional fixed-horizon tests waste samples when effects are obvious early—or fail to detect real improvements. The Soup repository solves this with a statistically rigorous sequential testing harness. This guide walks through the architecture, API usage, and CLI workflow backed by the actual source code in `MakazhanAlpamys/Soup`.

## Why Sequential A/B Testing Matters for Model Comparison

Fixed-sample t-tests break when you peek at results and stop early. The **mSPRT algorithm** in Soup guarantees that your `alpha` (false positive rate) and `beta` (false negative rate) stay valid even with optional stopping. This is implemented in [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) with decision boundaries computed from your error tolerances.

The framework supports three metrics natively: **`latency`**, **`judge_score`**, and **`retry_rate`**. These cover common production concerns—speed, quality, and reliability.

## Core Architecture Components

### MsprtConfig and MsprtVerdict Dataclasses

Configuration and results are type-safe dataclasses defined at [`src/soup_cli/utils/ab_test.py#L17-L27`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py#L17-L27) and [`src/soup_cli/utils/ab_test.py#L95-L104`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py#L95-L104):

```python
from soup_cli.utils.ab_test import MsprtConfig, MsprtVerdict

cfg = MsprtConfig(
    metric="latency",      # Must be: "latency", "judge_score", or "retry_rate"

    alpha=0.05,            # Max 5% false positive rate

    beta=0.20,             # Max 20% false negative rate

    effect_size=0.05       # Minimum detectable effect (standardized)

)

```

Validation enforces non-empty metric names under 32 characters through `validate_metric_name` at lines 46-65.

### The msprt_step Function

The core algorithm lives at [`src/soup_cli/utils/ab_test.py#L60-L145`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py#L60-L145). It:

1. Normalizes input samples to zero mean, unit variance
2. Computes pooled variance and standardized effect size (`z`)
3. Calculates log-likelihood-ratio using mixture prior
4. Compares LLR against upper/lower thresholds from `alpha` and `beta`
5. Returns a `MsprtVerdict` with `decision`, `log_likelihood_ratio`, and sample statistics

### File Ingestion with run_msprt

For batch processing, [`src/soup_cli/utils/ab_test.py#L56-L75`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py#L56-L75) streams JSONL files:

```python
from soup_cli.utils.ab_test import run_msprt

verdict = run_msprt(
    input_path="samples.jsonl",
    cfg=MsprtConfig(metric="latency", alpha=0.05, beta=0.20, effect_size=0.1)
)

```

Each JSONL line must contain `{"arm": "control" | "treatment", "<metric>": <float>}`.

## How to Run A/B Tests: Two Methods

### Method 1: Python API for Programmatic Control

Import directly when integrating into pipelines or notebooks:

```python
from soup_cli.utils.ab_test import MsprtConfig, msprt_step

# Production latency measurements (seconds)

control_latency = [0.245, 0.238, 0.251, 0.242, 0.239, 0.247]
treatment_latency = [0.198, 0.205, 0.192, 0.201, 0.196, 0.194]

cfg = MsprtConfig(
    metric="latency",
    alpha=0.05,        # 95% confidence

    beta=0.20,         # 80% power

    effect_size=0.10   # Detect 10% relative improvement

)

verdict = msprt_step(cfg, control=control_latency, treatment=treatment_latency)

print(f"Decision: {verdict.decision}")  # "reject_h0", "accept_h0", or "continue"

print(f"LLR: {verdict.log_likelihood_ratio:.4f}")
print(f"Samples: control={verdict.n_control}, treatment={verdict.n_treatment}")

```

Possible decisions:

- **`reject_h0`** – Treatment significantly outperforms control; safe to promote
- **`accept_h0`** – No detectable difference; retain control
- **`continue`** – Insufficient evidence; collect more samples

### Method 2: CLI for Operational Workflows

The `soup ab` command at [`src/soup_cli/commands/ab.py#L23-L43`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ab.py#L23-L43) provides a complete operational interface.

**Step 1:** Prepare your data file:

```bash
cat > model_latency.jsonl <<'EOF'
{"arm":"control","latency":0.245}
{"arm":"control","latency":0.238}
{"arm":"control","latency":0.251}
{"arm":"treatment","latency":0.198}
{"arm":"treatment","latency":0.205}
{"arm":"treatment","latency":0.192}
EOF

```

**Step 2:** Execute the test:

```bash
soup ab \
  --input model_latency.jsonl \
  --metric latency \
  --alpha 0.05 \
  --beta 0.20 \
  --effect-size 0.10

```

The output formats a verdict table with color-coded decisions.

### Adding Webhook Notifications

For production monitoring, attach Slack or Discord webhooks:

```bash
soup ab \
  --input judge_scores.jsonl \
  --metric judge_score \
  --alpha 0.01 \
  --beta 0.10 \
  --effect-size 0.05 \
  --slack-url "https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX"

```

Terminal decisions trigger POST requests with the full `MsprtVerdict` payload. Webhook logic is implemented in [[`src/soup_cli/commands/_webhook_cli.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/_webhook_cli.py)](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/_webhook_cli.py).

## Statistical Guarantees and Error Control

Unlike naive sequential testing, Soup's mSPRT maintains valid inference under **optional stopping**. You can:

- Check results after every sample
- Stop immediately when significance is reached
- Still trust your `alpha` and `beta` bounds

This property is documented in [[`docs/evaluation.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/evaluation.md)](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/evaluation.md#sequential-ab-harness-soup-ab) and derived from the mixture prior formulation in the LLR calculation.

## Complete File Reference

| File | Lines | Purpose |
|------|-------|---------|
| [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) | 17-27 | `MsprtVerdict` dataclass |
| [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) | 46-65 | `validate_metric_name` |
| [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) | 56-75 | `run_msprt` file processor |
| [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) | 60-145 | `msprt_step` core algorithm |
| [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) | 95-104 | `MsprtConfig` dataclass |
| [`src/soup_cli/commands/ab.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/ab.py) | 23-43 | CLI command `soup ab` |
| [`docs/evaluation.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/evaluation.md) | — | Statistical documentation |

## Summary

- **Soup implements mSPRT** for A/B testing model variants with guaranteed error control under optional stopping.
- **Three supported metrics**: `latency`, `judge_score`, and `retry_rate`—validated at initialization.
- **Two interfaces**: Direct Python API (`msprt_step`) for integration, CLI (`soup ab`) for operations.
- **Early stopping** reduces sample waste: decisions occur as soon as LLR crosses bounds, not at fixed horizons.
- **Production-ready** with webhook notifications and colored tabular output for monitoring.

## Frequently Asked Questions

### What metrics can I use for A/B testing in Soup?

Soup supports exactly three metrics: **`latency`** (response time), **`judge_score`** (quality evaluation), and **`retry_rate`** (reliability). The `validate_metric_name` function in [`src/soup_cli/utils/ab_test.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/ab_test.py) enforces this at runtime, rejecting invalid metrics with a clear error message.

### How does mSPRT prevent p-hacking compared to t-tests?

The **mSPRT computes decision thresholds from your `alpha` and `beta` upfront**, then compares the cumulative log-likelihood-ratio against these bounds. This mathematically guarantees that your error rates hold even if you peek repeatedly and stop as soon as you see significance. A t-test would inflate false positives under the same behavior.

### Can I extend Soup to support custom metrics?

Currently, metrics are hardcoded in `validate_metric_name`. To add metrics, you would modify the validation function and ensure downstream calculations in `msprt_step` handle the new data appropriately. The architecture is modular enough that metric-specific logic is isolated to normalization and variance pooling steps.

### When should I choose `continue` vs collecting more data?

The `continue` verdict means the LLR lies between the `accept_h0` and `reject_h0` bounds—in other words, evidence is ambiguous. You should gather more samples and re-run. The bounds tighten automatically as `n` increases, so you'll eventually reach a terminal decision if a real effect exists at your specified `effect_size`.