# How Switchyard's Soak-Test Methodology Evaluates Routing Overhead and Latency Under Load

> Switchyard's soak-test measures routing overhead and latency by comparing routed deployments to direct-backend baselines using deterministic traffic scenarios and tools like OHA and AIPerf.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-09-12

---

**Switchyard's soak-test framework measures routing overhead by comparing a routed Switchyard deployment against a direct-backend baseline using deterministic traffic scenarios, capturing delta latency and throughput metrics with tools like OHA and AIPerf.**

The NVIDIA-NeMo/Switchyard repository provides a comprehensive soak-test methodology designed to isolate and quantify the performance cost of routing decisions under sustained, high-concurrency workloads. By running identical request patterns through both a Switchyard release candidate and a direct-backend baseline, the framework calculates precise delta metrics that reveal the true latency and throughput impact of the routing layer.

## Core Components of the Soak-Test Framework

The evaluation relies on three tightly integrated components that work together to ensure accurate, reproducible measurements. These components are implemented across the `crates/switchyard-soak/` directory and orchestrated by Python scripts in `scripts/`.

### Traffic Generation with switchyard-soak

The Rust binary **`switchyard-soak`** serves as the primary load generator, executing predefined *scenarios* that exercise every routing path in the system. These scenarios include `short-interactive`, `decode-heavy`, `large-tool-catalog`, and resilience tests like `client-cancellation`.

Each scenario produces deterministic input files—specifically AIPerf `inputs-json` files and OHA payloads—that are replayed during benchmarking. The binary maintains **16 concurrent requests in flight** by default, configurable via the `--concurrency` flag, to simulate realistic pressure without overwhelming the system.

### Dual-Arm Benchmarking Architecture

The methodology simultaneously executes two parallel benchmark arms to establish a fair comparison baseline:

- **OHA (Open HTTP Accelerator)** measures raw HTTP request latency for short, non-interactive calls on every route.
- **AIPerf** replays exported streaming sessions, capturing **time-to-first-token (TTFT)**, **inter-token latency (ITL)**, **request throughput**, and **output-token throughput**.

The **direct-backend arm** (`--direct-base-url`/`--direct-model`) bypasses Switchyard entirely, hitting the same deployment that sits behind the router. The **routed arm** traverses the full Switchyard stack. Both tools run once per route with identical request payloads to ensure apples-to-apples comparison.

### Isolating Routing Overhead

The [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) script aggregates results from both arms and computes the critical delta metrics. By subtracting direct-backend measurements from routed measurements, the report yields **Δ latency** (positive values indicate extra routing time) and **Δ throughput** (negative values indicate reduced tokens per second).

The script outputs a comprehensive [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md) file alongside `report.csv`, [`report.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.json), and a visual `routing-overhead.svg` heat-map that highlights TTFT and token-throughput changes per route and workload. Additional telemetry includes classifier call counts, target-share percentages, and RSS/CPU growth to detect memory leaks or instability during the run.

## Methodological Guarantees for Accurate Measurement

To ensure that observed latency differences stem from routing logic rather than environmental variance, Switchyard's soak-test methodology enforces several strict constraints.

### Deterministic Backend and Scenario Catalog

The framework uses a local mock server (`switchyard-soak-mock`) that provides repeatable responses with fixed latency, isolating routing performance from model variance. The **scenario catalog** (documented in [`docs/operations/soak_test.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/operations/soak_test.md)) defines specific pressure angles—such as large tool catalogs to measure serialization overhead—that stress particular aspects of the routing logic.

### Fixed Concurrency and Statistical Confidence

Tests run with a fixed concurrency level (e.g., `--concurrency 100`) and request count (`--request-count 1000`) to guarantee that any latency change reflects routing overhead, not load variation. Multiple AIPerf runs configured via `--profile-runs 3` generate statistical confidence intervals for latency and throughput numbers.

### Failure Isolation and Health Monitoring

Resilience scenarios (failure-injection tests) run separately from performance benchmarks to prevent expected error rates from polluting throughput comparisons. During long-duration soaks, the framework monitors health endpoints (`/health`, `/metrics`) and tracks RSS growth, failing the test if memory exceeds thresholds like `--max-rss-growth-mib 512`.

## Running the Soak Tests

Execute the following commands to replicate the routing overhead evaluation locally:

```bash

# Build the server and soak tester (once per release)

cargo build --release -p switchyard-server
cargo build --release -p switchyard-soak

```

```bash

# Start the deterministic mock backend (optional)

cargo run --release -p switchyard-soak --example switchyard-soak-mock -- --port 8100 --latency-ms 40

```

```bash

# Compare routing algorithms under identical load

python3.12 scripts/benchmark_routing_algorithms.py \
  --base-url http://127.0.0.1:4000 \
  --direct-base-url http://127.0.0.1:8100 \
  --direct-model mock/weak \
  --model noop=switchyard/noop \
  --model passthrough=switchyard/passthrough \
  --model random=switchyard/random \
  --model llm_classifier=switchyard/classifier \
  --model stage_router=switchyard/stage \
  --concurrency 100 \
  --request-count 1000 \
  --scenario-set standard \
  --load-profile fixed \
  --profile-runs 3 \
  --backend-label "release model deployment"

```

This generates [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md), `report.csv`, [`report.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.json), and `routing-overhead.svg` in the current directory.

For extended stability validation, run a long-duration soak:

```bash
./target/release/switchyard-soak \
  --base-url http://127.0.0.1:4000 \
  --model RELEASE_MODEL_ID \
  --duration 48h \
  --concurrency 16 \
  --max-rss-growth-mib 512

```

The test writes timestamped results to `soak-results/` containing `intervals.csv`, [`summary.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/summary.json), and full execution logs.

## Summary

- Switchyard's soak-test methodology evaluates routing overhead by comparing delta metrics between a routed path and a direct-backend baseline using deterministic traffic.
- The `switchyard-soak` binary generates scenarios with fixed concurrency, while **OHA** measures HTTP latency and **AIPerf** captures streaming metrics like TTFT and ITL.
- The [`benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark_routing_algorithms.py) script isolates routing costs by computing **Δ latency** and **Δ throughput**, outputting results to [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md) and heat-map visualizations.
- Methodological rigor comes from deterministic mock backends, fixed load profiles, statistical confidence intervals via `--profile-runs`, and separate failure-injection testing.
- Long-duration soaks monitor RSS growth and health endpoints to verify stability under sustained load.

## Frequently Asked Questions

### How does Switchyard ensure that latency measurements reflect routing overhead rather than backend variability?

The framework uses the `switchyard-soak-mock` server to provide deterministic, repeatable responses with configurable fixed latency. By running the exact same request payloads through both the direct-backend arm and the routed arm, and by maintaining fixed concurrency levels, any variance in results can be attributed specifically to the routing layer's processing time.

### What metrics indicate a failure in the soak-test methodology?

According to [`docs/operations/soak_test.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/operations/soak_test.md), a test fails if **Δ latency** exceeds acceptable thresholds, **Δ throughput** drops significantly below baseline, or system health indicators degrade—specifically RSS growth exceeding the `--max-rss-growth-mib` limit or error rates spiking in non-failure-injection scenarios. The framework also validates that classifier call counts and target-share percentages match expected distributions.

### Can the soak-test methodology evaluate custom routing algorithms?

Yes. The [`benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark_routing_algorithms.py) script accepts multiple `--model` arguments that map arbitrary routing strategy identifiers to Switchyard model names. This allows direct comparison of custom implementations (such as `switchyard/stage` or `switchyard/classifier`) against baselines like `switchyard/passthrough` under identical load conditions defined by `--scenario-set`.

### What is the difference between OHA and AIPerf in the Switchyard soak test?

**OHA** (Open HTTP Accelerator) measures end-to-end HTTP request latency for short, non-streaming interactions, making it ideal for testing routing path overhead on simple requests. **AIPerf** specifically handles streaming sessions, measuring **time-to-first-token (TTFT)** and **inter-token latency (ITL)** to evaluate how routing decisions impact generative AI workloads where token throughput is critical.