How to Use A/B Testing to Compare Model Variants with the Soup Framework
The Soup framework implements sequential A/B testing for model variants using the Mixture Sequential Probability Ratio Test (mSPRT), enabling early stopping with guaranteed Type-I and Type-II error control.
A/B testing model variants is essential for production ML systems, but traditional fixed-horizon tests waste samples when effects are obvious early—or fail to detect real improvements. The Soup repository solves this with a statistically rigorous sequential testing harness. This guide walks through the architecture, API usage, and CLI workflow backed by the actual source code in MakazhanAlpamys/Soup.
Why Sequential A/B Testing Matters for Model Comparison
Fixed-sample t-tests break when you peek at results and stop early. The mSPRT algorithm in Soup guarantees that your alpha (false positive rate) and beta (false negative rate) stay valid even with optional stopping. This is implemented in src/soup_cli/utils/ab_test.py with decision boundaries computed from your error tolerances.
The framework supports three metrics natively: latency, judge_score, and retry_rate. These cover common production concerns—speed, quality, and reliability.
Core Architecture Components
MsprtConfig and MsprtVerdict Dataclasses
Configuration and results are type-safe dataclasses defined at src/soup_cli/utils/ab_test.py#L17-L27 and src/soup_cli/utils/ab_test.py#L95-L104:
from soup_cli.utils.ab_test import MsprtConfig, MsprtVerdict
cfg = MsprtConfig(
metric="latency", # Must be: "latency", "judge_score", or "retry_rate"
alpha=0.05, # Max 5% false positive rate
beta=0.20, # Max 20% false negative rate
effect_size=0.05 # Minimum detectable effect (standardized)
)
Validation enforces non-empty metric names under 32 characters through validate_metric_name at lines 46-65.
The msprt_step Function
The core algorithm lives at src/soup_cli/utils/ab_test.py#L60-L145. It:
- Normalizes input samples to zero mean, unit variance
- Computes pooled variance and standardized effect size (
z) - Calculates log-likelihood-ratio using mixture prior
- Compares LLR against upper/lower thresholds from
alphaandbeta - Returns a
MsprtVerdictwithdecision,log_likelihood_ratio, and sample statistics
File Ingestion with run_msprt
For batch processing, src/soup_cli/utils/ab_test.py#L56-L75 streams JSONL files:
from soup_cli.utils.ab_test import run_msprt
verdict = run_msprt(
input_path="samples.jsonl",
cfg=MsprtConfig(metric="latency", alpha=0.05, beta=0.20, effect_size=0.1)
)
Each JSONL line must contain {"arm": "control" | "treatment", "<metric>": <float>}.
How to Run A/B Tests: Two Methods
Method 1: Python API for Programmatic Control
Import directly when integrating into pipelines or notebooks:
from soup_cli.utils.ab_test import MsprtConfig, msprt_step
# Production latency measurements (seconds)
control_latency = [0.245, 0.238, 0.251, 0.242, 0.239, 0.247]
treatment_latency = [0.198, 0.205, 0.192, 0.201, 0.196, 0.194]
cfg = MsprtConfig(
metric="latency",
alpha=0.05, # 95% confidence
beta=0.20, # 80% power
effect_size=0.10 # Detect 10% relative improvement
)
verdict = msprt_step(cfg, control=control_latency, treatment=treatment_latency)
print(f"Decision: {verdict.decision}") # "reject_h0", "accept_h0", or "continue"
print(f"LLR: {verdict.log_likelihood_ratio:.4f}")
print(f"Samples: control={verdict.n_control}, treatment={verdict.n_treatment}")
Possible decisions:
reject_h0– Treatment significantly outperforms control; safe to promoteaccept_h0– No detectable difference; retain controlcontinue– Insufficient evidence; collect more samples
Method 2: CLI for Operational Workflows
The soup ab command at src/soup_cli/commands/ab.py#L23-L43 provides a complete operational interface.
Step 1: Prepare your data file:
cat > model_latency.jsonl <<'EOF'
{"arm":"control","latency":0.245}
{"arm":"control","latency":0.238}
{"arm":"control","latency":0.251}
{"arm":"treatment","latency":0.198}
{"arm":"treatment","latency":0.205}
{"arm":"treatment","latency":0.192}
EOF
Step 2: Execute the test:
soup ab \
--input model_latency.jsonl \
--metric latency \
--alpha 0.05 \
--beta 0.20 \
--effect-size 0.10
The output formats a verdict table with color-coded decisions.
Adding Webhook Notifications
For production monitoring, attach Slack or Discord webhooks:
soup ab \
--input judge_scores.jsonl \
--metric judge_score \
--alpha 0.01 \
--beta 0.10 \
--effect-size 0.05 \
--slack-url "https://hooks.slack.com/services/T00000000/B00000000/XXXXXXXXXXXXXXXXXXXXXXXX"
Terminal decisions trigger POST requests with the full MsprtVerdict payload. Webhook logic is implemented in [src/soup_cli/commands/_webhook_cli.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/_webhook_cli.py).
Statistical Guarantees and Error Control
Unlike naive sequential testing, Soup's mSPRT maintains valid inference under optional stopping. You can:
- Check results after every sample
- Stop immediately when significance is reached
- Still trust your
alphaandbetabounds
This property is documented in [docs/evaluation.md](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/evaluation.md#sequential-ab-harness-soup-ab) and derived from the mixture prior formulation in the LLR calculation.
Complete File Reference
| File | Lines | Purpose |
|---|---|---|
src/soup_cli/utils/ab_test.py |
17-27 | MsprtVerdict dataclass |
src/soup_cli/utils/ab_test.py |
46-65 | validate_metric_name |
src/soup_cli/utils/ab_test.py |
56-75 | run_msprt file processor |
src/soup_cli/utils/ab_test.py |
60-145 | msprt_step core algorithm |
src/soup_cli/utils/ab_test.py |
95-104 | MsprtConfig dataclass |
src/soup_cli/commands/ab.py |
23-43 | CLI command soup ab |
docs/evaluation.md |
— | Statistical documentation |
Summary
- Soup implements mSPRT for A/B testing model variants with guaranteed error control under optional stopping.
- Three supported metrics:
latency,judge_score, andretry_rate—validated at initialization. - Two interfaces: Direct Python API (
msprt_step) for integration, CLI (soup ab) for operations. - Early stopping reduces sample waste: decisions occur as soon as LLR crosses bounds, not at fixed horizons.
- Production-ready with webhook notifications and colored tabular output for monitoring.
Frequently Asked Questions
What metrics can I use for A/B testing in Soup?
Soup supports exactly three metrics: latency (response time), judge_score (quality evaluation), and retry_rate (reliability). The validate_metric_name function in src/soup_cli/utils/ab_test.py enforces this at runtime, rejecting invalid metrics with a clear error message.
How does mSPRT prevent p-hacking compared to t-tests?
The mSPRT computes decision thresholds from your alpha and beta upfront, then compares the cumulative log-likelihood-ratio against these bounds. This mathematically guarantees that your error rates hold even if you peek repeatedly and stop as soon as you see significance. A t-test would inflate false positives under the same behavior.
Can I extend Soup to support custom metrics?
Currently, metrics are hardcoded in validate_metric_name. To add metrics, you would modify the validation function and ensure downstream calculations in msprt_step handle the new data appropriately. The architecture is modular enough that metric-specific logic is isolated to normalization and variance pooling steps.
When should I choose continue vs collecting more data?
The continue verdict means the LLR lies between the accept_h0 and reject_h0 bounds—in other words, evidence is ambiguous. You should gather more samples and re-run. The bounds tighten automatically as n increases, so you'll eventually reach a terminal decision if a real effect exists at your specified effect_size.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →