# How Routing Overhead Is Measured During Switchyard Soak Testing

> Learn how Switchyard isolates routing overhead during soak testing. Discover how measuring latency delta reveals Switchyard specific processing time vs direct arm latency.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-09-13

---

**Switchyard isolates routing overhead by measuring the latency delta between a routed arm (requests processed through Switchyard) and a direct arm (requests sent straight to the backend), where the difference represents Switchyard-specific processing time.**

In the NVIDIA-NeMo/Switchyard repository, soak testing uses a controlled dual-arm experiment to separate **routing overhead** from total request latency. Unlike raw end-to-end latency—which aggregates backend processing with routing decisions—this measurement isolates the specific computational cost added by Switchyard's routing algorithm, scoring mechanisms, and target selection logic.

## The Dual-Arm Benchmark Methodology

The measurement strategy relies on executing two parallel runs of identical workloads. This approach isolates variables by using the same backend deployment and model configurations in both arms, ensuring that any latency differences stem solely from Switchyard's routing layer.

### The Routed Arm

The routed arm sends requests through Switchyard's full processing pipeline. In [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py), this corresponds to the standard benchmark configuration where requests hit the Switchyard endpoint and the routing algorithm selects the target deployment. The script captures **end-to-end latency** metrics including time-to-first-token (TTFT) and total request duration as the request travels through the routing, scoring, and selection layers.

### The Direct Arm

The direct arm bypasses Switchyard entirely, sending identical requests straight to the backend deployment. As implemented in [`benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark_routing_algorithms.py), you configure this using the `--direct-base-url` and `--direct-model` options. This arm uses `oha` (a load testing tool) to hit the backend directly, establishing a baseline latency that includes only backend processing time without any Switchyard overhead.

### Delta Calculation Logic

After both arms complete, Switchyard computes the difference between routed and direct latencies. According to [`docs/operations/soak_test.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/operations/soak_test.md) (lines 97-103), the calculation follows these principles:

- **Positive TTFT or request latency delta** → The routed request took longer than the direct request; this extra time is attributed to **routing overhead**.
- **Negative token-throughput delta** → The routed path processed fewer tokens per second than the direct path; this reduction also counts as routing overhead.

Mathematically: `Routing Overhead = Routed Latency - Direct Latency`.

## Routing Overhead vs. End-to-End Latency

Understanding the distinction between these metrics is critical for performance analysis.

**Routing overhead** represents only the incremental processing time added by Switchyard's routing layer. It excludes backend computation time and focuses exclusively on the cost of routing decisions, load balancing, and target selection. This metric answers the question: "How much slower is routing through Switchyard compared to hitting the backend directly?"

**End-to-end latency** encompasses the total time from request receipt to final token generation, including both backend processing *and* routing overhead. As noted in the soak test documentation, this is the sum of direct latency plus routing overhead: `End-to-End Latency = Direct Latency + Routing Overhead`.

When examining [`docs/operations/soak_test.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/operations/soak_test.md) (lines 24-27), you'll find that the soak test explicitly strips out backend processing time to isolate the overhead component, whereas raw end-to-end latency aggregates both components without separation.

## Implementing the Measurement Pipeline

To reproduce these measurements in your own Switchyard deployment, execute the following commands.

First, build the necessary binaries:

```bash
cargo build --release -p switchyard-server
cargo build --release -p switchyard-soak

```

Run the soak tester to generate scenario data:

```bash
./target/release/switchyard-soak \
    --base-url http://127.0.0.1:4000 \
    --model switchyard/general \
    --duration 5m \
    --concurrency 4

```

Execute the dual-arm benchmark to compute overhead:

```bash
python3.12 scripts/benchmark_routing_algorithms.py \
    --base-url http://127.0.0.1:4000 \
    --direct-base-url http://127.0.0.1:8100 \
    --direct-model mock/weak \
    --model noop=switchyard/noop \
    --model passthrough=switchyard/passthrough \
    --model random=switchyard/random \
    --model llm_classifier=switchyard/classifier \
    --model stage_router=switchyard/stage \
    --concurrency 100 \
    --request-count 1000 \
    --scenario-set standard \
    --load-profile fixed \
    --profile-runs 3 \
    --backend-label "release model deployment"

```

The benchmark script performs three distinct operations:

1. Exports scenario payloads using `switchyard-soak`.
2. Replays them with **AIPerf** for the routed arm and **oha** for the direct arm.
3. Computes per-route latency differences and writes results to [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md), `report.csv`, and [`report.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.json), while generating `routing-overhead.svg`.

## Analyzing the Results

The soak test generates structured reports that visualize routing overhead separately from total latency. In [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md), you'll encounter rows formatted as:

```

| Route | Workload | TTFT Δ (ms) | Tokens-throughput Δ (tps) |
|-------|----------|-------------|---------------------------|
| random | short-interactive | +12.4 | -3.2 |

```

Interpret these values as follows:

- `+12.4 ms` indicates the routed request incurred 12.4 milliseconds of **routing overhead** compared to the direct request.
- `-3.2 tps` shows the routed path processed 3.2 fewer tokens per second, representing throughput degradation attributable to routing logic.

These deltas are not raw end-to-end latencies; they represent the isolated cost of Switchyard's processing after subtracting the direct baseline. The visualization in `routing-overhead.svg` plots these differentials across various routing algorithms and workloads.

The measurement pipeline relies on several key source files:

- [`crates/switchyard-soak/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/src/lib.rs) (lines 365-371) handles scenario export and timestamp collection for latency calculations.
- [`crates/switchyard-soak/src/stats.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/src/stats.rs) (lines 84-90) defines in-memory structures for tracking per-run statistics that feed into overhead calculations.

## Summary

- **Routing overhead** is calculated as the latency difference between routed requests (through Switchyard) and direct requests (straight to backend), isolating Switchyard-specific processing time.
- **End-to-end latency** combines backend processing time with routing overhead, representing the total observable latency when using Switchyard.
- The dual-arm methodology uses [`benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark_routing_algorithms.py) with `--direct-base-url` and `--direct-model` options to establish baselines.
- Positive TTFT deltas and negative throughput deltas in the soak test reports indicate routing overhead.
- Results are emitted in [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md), `report.csv`, and visualized in `routing-overhead.svg` for analysis.

## Frequently Asked Questions

### How does Switchyard ensure the direct arm provides an accurate baseline?

The direct arm uses identical backend deployments, models, and response configurations as the routed arm, but bypasses Switchyard entirely using the `--direct-base-url` parameter. Because both arms process the same scenario payloads generated by `switchyard-soak`, any latency differences can only originate from Switchyard's routing layer, not from backend variability.

### Why is time-to-first-token (TTFT) used instead of total request duration?

TTFT isolates the initial routing decision and connection overhead, which are the primary contributions of the routing layer. While total request duration is also measured, TTFT specifically captures the latency impact of target selection and scoring algorithms before token generation begins, making it a sensitive indicator of routing overhead.

### Can routing overhead measurements be negative?

Yes. If Switchyard's routing algorithm selects a faster backend instance or optimizes connection pooling better than the direct arm's static configuration, the routed arm may show lower latency than the direct arm. This negative delta indicates that Switchyard's intelligent routing actually reduced latency compared to the direct baseline.

### Where are the raw latency metrics stored during soak testing?

Raw metrics are stored in-memory using structures defined in [`crates/switchyard-soak/src/stats.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/src/stats.rs) (lines 84-90) during execution, then serialized to [`report.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.json), `report.csv`, and [`report.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/report.md) upon completion. The timestamp data used for latency calculations is collected in [`crates/switchyard-soak/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/src/lib.rs) (lines 365-371) and processed by [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) to generate the final overhead differentials.