How to Use A/B Testing to Compare Model Variants with the Soup Framework

The Soup framework implements sequential A/B testing for model variants using the Mixture Sequential Probability Ratio Test (mSPRT), enabling early stopping with guaranteed Type-I and Type-II error control.

A/B testing model variants is essential for production ML systems, but traditional fixed-horizon tests waste samples when effects are obvious early—or fail to detect real improvements. The Soup repository solves this with a statistically rigorous sequential testing harness. This guide walks through the architecture, API usage, and CLI workflow backed by the actual source code in MakazhanAlpamys/Soup.

Why Sequential A/B Testing Matters for Model Comparison

Fixed-sample t-tests break when you peek at results and stop early. The mSPRT algorithm in Soup guarantees that your alpha (false positive rate) and beta (false negative rate) stay valid even with optional stopping. This is implemented in src/soup_cli/utils/ab_test.py with decision boundaries computed from your error tolerances.

The framework supports three metrics natively: latency, judge_score, and retry_rate. These cover common production concerns—speed, quality, and reliability.

Core Architecture Components

MsprtConfig and MsprtVerdict Dataclasses

Configuration and results are type-safe dataclasses defined at src/soup_cli/utils/ab_test.py#L17-L27 and src/soup_cli/utils/ab_test.py#L95-L104:

from soup_cli.utils.ab_test import MsprtConfig, MsprtVerdict

cfg = MsprtConfig(
    metric="latency",      # Must be: "latency", "judge_score", or "retry_rate"

    alpha=0.05,            # Max 5% false positive rate

    beta=0.20,             # Max 20% false negative rate

    effect_size=0.05       # Minimum detectable effect (standardized)

)

Validation enforces non-empty metric names under 32 characters through validate_metric_name at lines 46-65.

The msprt_step Function

The core algorithm lives at src/soup_cli/utils/ab_test.py#L60-L145. It:

  1. Normalizes input samples to zero mean, unit variance
  2. Computes pooled variance and standardized effect size (z)
  3. Calculates log-likelihood-ratio using mixture prior
  4. Compares LLR against upper/lower thresholds from alpha and beta
  5. Returns a MsprtVerdict with decision, log_likelihood_ratio, and sample statistics

File Ingestion with run_msprt

For batch processing, src/soup_cli/utils/ab_test.py#L56-L75 streams JSONL files:

from soup_cli.utils.ab_test import run_msprt

verdict = run_msprt(
    input_path="samples.jsonl",
    cfg=MsprtConfig(metric="latency", alpha=0.05, beta=0.20, effect_size=0.1)
)

Each JSONL line must contain {"arm": "control" | "treatment", "<metric>": <float>}.

How to Run A/B Tests: Two Methods

Method 1: Python API for Programmatic Control

Import directly when integrating into pipelines or notebooks:

from soup_cli.utils.ab_test import MsprtConfig, msprt_step

# Production latency measurements (seconds)

control_latency = [0.245, 0.238, 0.251, 0.242, 0.239, 0.247]
treatment_latency = [0.198, 0.205, 0.192, 0.201, 0.196, 0.194]

cfg = MsprtConfig(
    metric="latency",
    alpha=0.05,        # 95% confidence

    beta=0.20,         # 80% power

    effect_size=0.10   # Detect 10% relative improvement

)

verdict = msprt_step(cfg, control=control_latency, treatment=treatment_latency)

print(f"Decision: {verdict.decision}")  # "reject_h0", "accept_h0", or "continue"

print(f"LLR: {verdict.log_likelihood_ratio:.4f}")
print(f"Samples: control={verdict.n_control}, treatment={verdict.n_treatment}")

Possible decisions:

  • reject_h0 – Treatment significantly outperforms control; safe to promote
  • accept_h0 – No detectable difference; retain control
  • continue – Insufficient evidence; collect more samples

Method 2: CLI for Operational Workflows

The soup ab command at src/soup_cli/commands/ab.py#L23-L43 provides a complete operational interface.

Step 1: Prepare your data file:

cat > model_latency.jsonl <<'EOF'
{"arm":"control","latency":0.245}
{"arm":"control","latency":0.238}
{"arm":"control","latency":0.251}
{"arm":"treatment","latency":0.198}
{"arm":"treatment","latency":0.205}
{"arm":"treatment","latency":0.192}
EOF

Step 2: Execute the test:

soup ab \
  --input model_latency.jsonl \
  --metric latency \
  --alpha 0.05 \
  --beta 0.20 \
  --effect-size 0.10

The output formats a verdict table with color-coded decisions.

Adding Webhook Notifications

For production monitoring, attach Slack or Discord webhooks:

soup ab \
  --input judge_scores.jsonl \
  --metric judge_score \
  --alpha 0.01 \
  --beta 0.10 \
  --effect-size 0.05 \
  --slack-url "https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX"

Terminal decisions trigger POST requests with the full MsprtVerdict payload. Webhook logic is implemented in [src/soup_cli/commands/_webhook_cli.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/_webhook_cli.py).

Statistical Guarantees and Error Control

Unlike naive sequential testing, Soup's mSPRT maintains valid inference under optional stopping. You can:

  • Check results after every sample
  • Stop immediately when significance is reached
  • Still trust your alpha and beta bounds

This property is documented in [docs/evaluation.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/evaluation.md#sequential-ab-harness-soup-ab) and derived from the mixture prior formulation in the LLR calculation.

Complete File Reference

File Lines Purpose
src/soup_cli/utils/ab_test.py 17-27 MsprtVerdict dataclass
src/soup_cli/utils/ab_test.py 46-65 validate_metric_name
src/soup_cli/utils/ab_test.py 56-75 run_msprt file processor
src/soup_cli/utils/ab_test.py 60-145 msprt_step core algorithm
src/soup_cli/utils/ab_test.py 95-104 MsprtConfig dataclass
src/soup_cli/commands/ab.py 23-43 CLI command soup ab
docs/evaluation.md — Statistical documentation

Summary

  • Soup implements mSPRT for A/B testing model variants with guaranteed error control under optional stopping.
  • Three supported metrics: latency, judge_score, and retry_rate—validated at initialization.
  • Two interfaces: Direct Python API (msprt_step) for integration, CLI (soup ab) for operations.
  • Early stopping reduces sample waste: decisions occur as soon as LLR crosses bounds, not at fixed horizons.
  • Production-ready with webhook notifications and colored tabular output for monitoring.

Frequently Asked Questions

What metrics can I use for A/B testing in Soup?

Soup supports exactly three metrics: latency (response time), judge_score (quality evaluation), and retry_rate (reliability). The validate_metric_name function in src/soup_cli/utils/ab_test.py enforces this at runtime, rejecting invalid metrics with a clear error message.

How does mSPRT prevent p-hacking compared to t-tests?

The mSPRT computes decision thresholds from your alpha and beta upfront, then compares the cumulative log-likelihood-ratio against these bounds. This mathematically guarantees that your error rates hold even if you peek repeatedly and stop as soon as you see significance. A t-test would inflate false positives under the same behavior.

Can I extend Soup to support custom metrics?

Currently, metrics are hardcoded in validate_metric_name. To add metrics, you would modify the validation function and ensure downstream calculations in msprt_step handle the new data appropriately. The architecture is modular enough that metric-specific logic is isolated to normalization and variance pooling steps.

When should I choose continue vs collecting more data?

The continue verdict means the LLR lies between the accept_h0 and reject_h0 bounds—in other words, evidence is ambiguous. You should gather more samples and re-run. The bounds tighten automatically as n increases, so you'll eventually reach a terminal decision if a real effect exists at your specified effect_size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →