How Routing Overhead Is Measured During Switchyard Soak Testing

Switchyard isolates routing overhead by measuring the latency delta between a routed arm (requests processed through Switchyard) and a direct arm (requests sent straight to the backend), where the difference represents Switchyard-specific processing time.

In the NVIDIA-NeMo/Switchyard repository, soak testing uses a controlled dual-arm experiment to separate routing overhead from total request latency. Unlike raw end-to-end latency—which aggregates backend processing with routing decisions—this measurement isolates the specific computational cost added by Switchyard's routing algorithm, scoring mechanisms, and target selection logic.

The Dual-Arm Benchmark Methodology

The measurement strategy relies on executing two parallel runs of identical workloads. This approach isolates variables by using the same backend deployment and model configurations in both arms, ensuring that any latency differences stem solely from Switchyard's routing layer.

The Routed Arm

The routed arm sends requests through Switchyard's full processing pipeline. In scripts/benchmark_routing_algorithms.py, this corresponds to the standard benchmark configuration where requests hit the Switchyard endpoint and the routing algorithm selects the target deployment. The script captures end-to-end latency metrics including time-to-first-token (TTFT) and total request duration as the request travels through the routing, scoring, and selection layers.

The Direct Arm

The direct arm bypasses Switchyard entirely, sending identical requests straight to the backend deployment. As implemented in benchmark_routing_algorithms.py, you configure this using the --direct-base-url and --direct-model options. This arm uses oha (a load testing tool) to hit the backend directly, establishing a baseline latency that includes only backend processing time without any Switchyard overhead.

Delta Calculation Logic

After both arms complete, Switchyard computes the difference between routed and direct latencies. According to docs/operations/soak_test.md (lines 97-103), the calculation follows these principles:

  • Positive TTFT or request latency delta → The routed request took longer than the direct request; this extra time is attributed to routing overhead.
  • Negative token-throughput delta → The routed path processed fewer tokens per second than the direct path; this reduction also counts as routing overhead.

Mathematically: Routing Overhead = Routed Latency - Direct Latency.

Routing Overhead vs. End-to-End Latency

Understanding the distinction between these metrics is critical for performance analysis.

Routing overhead represents only the incremental processing time added by Switchyard's routing layer. It excludes backend computation time and focuses exclusively on the cost of routing decisions, load balancing, and target selection. This metric answers the question: "How much slower is routing through Switchyard compared to hitting the backend directly?"

End-to-end latency encompasses the total time from request receipt to final token generation, including both backend processing and routing overhead. As noted in the soak test documentation, this is the sum of direct latency plus routing overhead: End-to-End Latency = Direct Latency + Routing Overhead.

When examining docs/operations/soak_test.md (lines 24-27), you'll find that the soak test explicitly strips out backend processing time to isolate the overhead component, whereas raw end-to-end latency aggregates both components without separation.

Implementing the Measurement Pipeline

To reproduce these measurements in your own Switchyard deployment, execute the following commands.

First, build the necessary binaries:

cargo build --release -p switchyard-server
cargo build --release -p switchyard-soak

Run the soak tester to generate scenario data:

./target/release/switchyard-soak \
    --base-url http://127.0.0.1:4000 \
    --model switchyard/general \
    --duration 5m \
    --concurrency 4

Execute the dual-arm benchmark to compute overhead:

python3.12 scripts/benchmark_routing_algorithms.py \
    --base-url http://127.0.0.1:4000 \
    --direct-base-url http://127.0.0.1:8100 \
    --direct-model mock/weak \
    --model noop=switchyard/noop \
    --model passthrough=switchyard/passthrough \
    --model random=switchyard/random \
    --model llm_classifier=switchyard/classifier \
    --model stage_router=switchyard/stage \
    --concurrency 100 \
    --request-count 1000 \
    --scenario-set standard \
    --load-profile fixed \
    --profile-runs 3 \
    --backend-label "release model deployment"

The benchmark script performs three distinct operations:

  1. Exports scenario payloads using switchyard-soak.
  2. Replays them with AIPerf for the routed arm and oha for the direct arm.
  3. Computes per-route latency differences and writes results to report.md, report.csv, and report.json, while generating routing-overhead.svg.

Analyzing the Results

The soak test generates structured reports that visualize routing overhead separately from total latency. In report.md, you'll encounter rows formatted as:


| Route | Workload | TTFT Δ (ms) | Tokens-throughput Δ (tps) |
|-------|----------|-------------|---------------------------|
| random | short-interactive | +12.4 | -3.2 |

Interpret these values as follows:

  • +12.4 ms indicates the routed request incurred 12.4 milliseconds of routing overhead compared to the direct request.
  • -3.2 tps shows the routed path processed 3.2 fewer tokens per second, representing throughput degradation attributable to routing logic.

These deltas are not raw end-to-end latencies; they represent the isolated cost of Switchyard's processing after subtracting the direct baseline. The visualization in routing-overhead.svg plots these differentials across various routing algorithms and workloads.

The measurement pipeline relies on several key source files:

Summary

  • Routing overhead is calculated as the latency difference between routed requests (through Switchyard) and direct requests (straight to backend), isolating Switchyard-specific processing time.
  • End-to-end latency combines backend processing time with routing overhead, representing the total observable latency when using Switchyard.
  • The dual-arm methodology uses benchmark_routing_algorithms.py with --direct-base-url and --direct-model options to establish baselines.
  • Positive TTFT deltas and negative throughput deltas in the soak test reports indicate routing overhead.
  • Results are emitted in report.md, report.csv, and visualized in routing-overhead.svg for analysis.

Frequently Asked Questions

How does Switchyard ensure the direct arm provides an accurate baseline?

The direct arm uses identical backend deployments, models, and response configurations as the routed arm, but bypasses Switchyard entirely using the --direct-base-url parameter. Because both arms process the same scenario payloads generated by switchyard-soak, any latency differences can only originate from Switchyard's routing layer, not from backend variability.

Why is time-to-first-token (TTFT) used instead of total request duration?

TTFT isolates the initial routing decision and connection overhead, which are the primary contributions of the routing layer. While total request duration is also measured, TTFT specifically captures the latency impact of target selection and scoring algorithms before token generation begins, making it a sensitive indicator of routing overhead.

Can routing overhead measurements be negative?

Yes. If Switchyard's routing algorithm selects a faster backend instance or optimizes connection pooling better than the direct arm's static configuration, the routed arm may show lower latency than the direct arm. This negative delta indicates that Switchyard's intelligent routing actually reduced latency compared to the direct baseline.

Where are the raw latency metrics stored during soak testing?

Raw metrics are stored in-memory using structures defined in crates/switchyard-soak/src/stats.rs (lines 84-90) during execution, then serialized to report.json, report.csv, and report.md upon completion. The timestamp data used for latency calculations is collected in crates/switchyard-soak/src/lib.rs (lines 365-371) and processed by scripts/benchmark_routing_algorithms.py to generate the final overhead differentials.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →