How Switchyard's Soak-Test Methodology Evaluates Routing Overhead and Latency Under Load

Switchyard's soak-test framework measures routing overhead by comparing a routed Switchyard deployment against a direct-backend baseline using deterministic traffic scenarios, capturing delta latency and throughput metrics with tools like OHA and AIPerf.

The NVIDIA-NeMo/Switchyard repository provides a comprehensive soak-test methodology designed to isolate and quantify the performance cost of routing decisions under sustained, high-concurrency workloads. By running identical request patterns through both a Switchyard release candidate and a direct-backend baseline, the framework calculates precise delta metrics that reveal the true latency and throughput impact of the routing layer.

Core Components of the Soak-Test Framework

The evaluation relies on three tightly integrated components that work together to ensure accurate, reproducible measurements. These components are implemented across the crates/switchyard-soak/ directory and orchestrated by Python scripts in scripts/.

Traffic Generation with switchyard-soak

The Rust binary switchyard-soak serves as the primary load generator, executing predefined scenarios that exercise every routing path in the system. These scenarios include short-interactive, decode-heavy, large-tool-catalog, and resilience tests like client-cancellation.

Each scenario produces deterministic input files—specifically AIPerf inputs-json files and OHA payloads—that are replayed during benchmarking. The binary maintains 16 concurrent requests in flight by default, configurable via the --concurrency flag, to simulate realistic pressure without overwhelming the system.

Dual-Arm Benchmarking Architecture

The methodology simultaneously executes two parallel benchmark arms to establish a fair comparison baseline:

  • OHA (Open HTTP Accelerator) measures raw HTTP request latency for short, non-interactive calls on every route.
  • AIPerf replays exported streaming sessions, capturing time-to-first-token (TTFT), inter-token latency (ITL), request throughput, and output-token throughput.

The direct-backend arm (--direct-base-url/--direct-model) bypasses Switchyard entirely, hitting the same deployment that sits behind the router. The routed arm traverses the full Switchyard stack. Both tools run once per route with identical request payloads to ensure apples-to-apples comparison.

Isolating Routing Overhead

The scripts/benchmark_routing_algorithms.py script aggregates results from both arms and computes the critical delta metrics. By subtracting direct-backend measurements from routed measurements, the report yields Δ latency (positive values indicate extra routing time) and Δ throughput (negative values indicate reduced tokens per second).

The script outputs a comprehensive report.md file alongside report.csv, report.json, and a visual routing-overhead.svg heat-map that highlights TTFT and token-throughput changes per route and workload. Additional telemetry includes classifier call counts, target-share percentages, and RSS/CPU growth to detect memory leaks or instability during the run.

Methodological Guarantees for Accurate Measurement

To ensure that observed latency differences stem from routing logic rather than environmental variance, Switchyard's soak-test methodology enforces several strict constraints.

Deterministic Backend and Scenario Catalog

The framework uses a local mock server (switchyard-soak-mock) that provides repeatable responses with fixed latency, isolating routing performance from model variance. The scenario catalog (documented in docs/operations/soak_test.md) defines specific pressure angles—such as large tool catalogs to measure serialization overhead—that stress particular aspects of the routing logic.

Fixed Concurrency and Statistical Confidence

Tests run with a fixed concurrency level (e.g., --concurrency 100) and request count (--request-count 1000) to guarantee that any latency change reflects routing overhead, not load variation. Multiple AIPerf runs configured via --profile-runs 3 generate statistical confidence intervals for latency and throughput numbers.

Failure Isolation and Health Monitoring

Resilience scenarios (failure-injection tests) run separately from performance benchmarks to prevent expected error rates from polluting throughput comparisons. During long-duration soaks, the framework monitors health endpoints (/health, /metrics) and tracks RSS growth, failing the test if memory exceeds thresholds like --max-rss-growth-mib 512.

Running the Soak Tests

Execute the following commands to replicate the routing overhead evaluation locally:


# Build the server and soak tester (once per release)

cargo build --release -p switchyard-server
cargo build --release -p switchyard-soak

# Start the deterministic mock backend (optional)

cargo run --release -p switchyard-soak --example switchyard-soak-mock -- --port 8100 --latency-ms 40

# Compare routing algorithms under identical load

python3.12 scripts/benchmark_routing_algorithms.py \
  --base-url http://127.0.0.1:4000 \
  --direct-base-url http://127.0.0.1:8100 \
  --direct-model mock/weak \
  --model noop=switchyard/noop \
  --model passthrough=switchyard/passthrough \
  --model random=switchyard/random \
  --model llm_classifier=switchyard/classifier \
  --model stage_router=switchyard/stage \
  --concurrency 100 \
  --request-count 1000 \
  --scenario-set standard \
  --load-profile fixed \
  --profile-runs 3 \
  --backend-label "release model deployment"

This generates report.md, report.csv, report.json, and routing-overhead.svg in the current directory.

For extended stability validation, run a long-duration soak:

./target/release/switchyard-soak \
  --base-url http://127.0.0.1:4000 \
  --model RELEASE_MODEL_ID \
  --duration 48h \
  --concurrency 16 \
  --max-rss-growth-mib 512

The test writes timestamped results to soak-results/ containing intervals.csv, summary.json, and full execution logs.

Summary

  • Switchyard's soak-test methodology evaluates routing overhead by comparing delta metrics between a routed path and a direct-backend baseline using deterministic traffic.
  • The switchyard-soak binary generates scenarios with fixed concurrency, while OHA measures HTTP latency and AIPerf captures streaming metrics like TTFT and ITL.
  • The benchmark_routing_algorithms.py script isolates routing costs by computing Δ latency and Δ throughput, outputting results to report.md and heat-map visualizations.
  • Methodological rigor comes from deterministic mock backends, fixed load profiles, statistical confidence intervals via --profile-runs, and separate failure-injection testing.
  • Long-duration soaks monitor RSS growth and health endpoints to verify stability under sustained load.

Frequently Asked Questions

How does Switchyard ensure that latency measurements reflect routing overhead rather than backend variability?

The framework uses the switchyard-soak-mock server to provide deterministic, repeatable responses with configurable fixed latency. By running the exact same request payloads through both the direct-backend arm and the routed arm, and by maintaining fixed concurrency levels, any variance in results can be attributed specifically to the routing layer's processing time.

What metrics indicate a failure in the soak-test methodology?

According to docs/operations/soak_test.md, a test fails if Δ latency exceeds acceptable thresholds, Δ throughput drops significantly below baseline, or system health indicators degrade—specifically RSS growth exceeding the --max-rss-growth-mib limit or error rates spiking in non-failure-injection scenarios. The framework also validates that classifier call counts and target-share percentages match expected distributions.

Can the soak-test methodology evaluate custom routing algorithms?

Yes. The benchmark_routing_algorithms.py script accepts multiple --model arguments that map arbitrary routing strategy identifiers to Switchyard model names. This allows direct comparison of custom implementations (such as switchyard/stage or switchyard/classifier) against baselines like switchyard/passthrough under identical load conditions defined by --scenario-set.

What is the difference between OHA and AIPerf in the Switchyard soak test?

OHA (Open HTTP Accelerator) measures end-to-end HTTP request latency for short, non-streaming interactions, making it ideal for testing routing path overhead on simple requests. AIPerf specifically handles streaming sessions, measuring time-to-first-token (TTFT) and inter-token latency (ITL) to evaluate how routing decisions impact generative AI workloads where token throughput is critical.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →