# Switchyard Performance Considerations: Latency, Throughput, and Scaling Explained

> Understand Switchyard performance considerations: latency, throughput, and scaling. Discover minimal overhead and linear scaling for LLM requests.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-08-23

---

**TLDR: Switchyard adds only a few milliseconds of routing overhead per LLM request, scales linearly with concurrency, preserves backend error semantics, and ships with a built-in benchmark suite in [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) to measure every performance dimension.**

Switchyard, an open-source orchestration layer from **NVIDIA-NeMo/Switchyard**, sits between client applications and large-language-model backends, classifying and forwarding each request. While such a routing layer can introduce overhead, the project's design and benchmark tooling are built to keep that impact minimal. This article breaks down the key performance considerations for Switchyard, the metrics that matter, and exactly how to measure them using the repository's own benchmark suite.

## The Performance Dimensions Switchyard Tracks

Switchyard's benchmark suite categorizes performance impact across five measurable dimensions. Each maps to a concrete metric collected during load testing:

| Dimension | What it measures | Typical metrics |
|-----------|------------------|-----------------|
| **Routing latency** | Time spent classifying and forwarding each request | `routing_overhead_avg_ms` |
| **Throughput** | Requests or tokens processed per second | `oha_requests_per_second`, `aiperf_requests_per_second`, `aiperf_output_tokens_per_second` |
| **Error handling** | Ability to keep error rates low while routing | `aiperf_error_rate`, `expected_error_rate_min/max` |
| **Concurrency scaling** | Behavior under high parallel load | `concurrency`, `request_rate` |
| **Resource usage** | Memory and request-size limits for materialized traffic | `MAX_MATERIALIZED_REQUESTS`, `MAX_MATERIALIZED_BYTES` |

The core benchmark code that collects these metrics lives in **[`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py)**, and the accompanying test suite ([`tests/test_routing_performance_report.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_routing_performance_report.py)) validates that the reporting logic computes and aggregates each value correctly.

## The Five Main Ways Switchyard Affects Performance

### 1. Routing Overhead

Every request passes through a classifier before reaching the chosen backend. That classification cost is captured in the `routing_overhead_avg_ms` field, which is defined by the `RoutingDelta` dataclass in the benchmark script. In practice, benchmarks show overheads of **a few milliseconds per request** — negligible compared to the inference latency of most LLM backends.

### 2. Concurrent Request Handling

Switchyard is fully asynchronous and sustains high concurrency. The `concurrency` parameter in the benchmark harness drives both `oha` and AIPerf to verify that Switchyard scales linearly, up to the limits of the underlying HTTP server and the Python event loop.

### 3. Token Throughput

Because Switchyard forwards the payload unmodified, token throughput (`aiperf_output_tokens_per_second`) is dominated by the backend model, not the router. The benchmark measures the delta between a direct backend call and a routed call to confirm the routing layer adds **no extra token-processing bottleneck**.

### 4. Error Propagation

The routing logic must preserve backend error semantics exactly. The test suite compares error counts and rates from a direct backend (`compare_to_direct_backend`) against those observed through Switchyard, ensuring error amplification stays within a tight bound, as enforced by the `expected_error_rate_min/max` fields.

### 5. Request Size Handling

The `MAX_MATERIALIZED_REQUESTS` constant caps the number of in-memory requests, and `MAX_MATERIALIZED_BYTES` caps the total byte size. These two limits prevent out-of-memory conditions while still supplying realistic traffic patterns for benchmarks.

## Running the Switchyard Benchmark

The benchmark suite ships with the repository in the `scripts/` directory. Install development dependencies and launch the server, then run the benchmark:

```bash

# Install development dependencies

uv sync --group dev

# Launch a Switchyard server (example config)

switchyard-server --config routes.toml --port 4000

# Run the benchmark (replace <BASE_URL> with your server address)

python -m scripts.benchmark_routing_algorithms \
    --base-url http://127.0.0.1:4000 \
    --concurrency 100 \
    --request-count 1000 \
    --output-dir ./benchmark-results

```

The script then:

1. Exports scenarios from the native Rust "soak" tool.
2. Executes a short baseline with `oha` and a full load test with AIPerf.
3. Computes routing overhead, throughput, and latency deltas.
4. Emits a markdown report (via the `write_report` function) and an optional CSV for downstream analysis.

## Interpreting the Benchmark Report

A typical report row from the generated markdown looks like this:

| algorithm | model | scenario | oha_rps | oha_p50 ms | aiperf_rps | aiperf_p50 ms | routing_overhead ms |
|-----------|-------|----------|---------|------------|------------|---------------|---------------------|
| `random`  | `switchyard/random` | `short-interactive` | 120.5 | 0.012 | 115.3 | 0.018 | 2.4 |

Key columns to focus on:

- **`routing_overhead_ms`** – The extra latency introduced by Switchyard's classifier. Values in the low single-digit milliseconds are typical and expected.
- **`oha_rps` vs `aiperf_rps`** – Baseline versus routed requests per-second. A small drop indicates acceptable overhead; a large gap suggests a tuning opportunity.
- **Error columns** – `aiperf_error_rate` should match the direct backend's rate, and the test suite enforces this equality.

## Key Source Files for Performance

| File | Role |
|------|------|
| [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) | Main benchmark driver that materializes traffic, runs `oha` and AIPerf, and writes performance reports |
| [`tests/test_routing_performance_report.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_routing_performance_report.py) | Unit test validating the comparison of routed versus direct backend results and ensuring correct aggregation in the report |
| [`scripts/routing_overhead_plot.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/routing_overhead_plot.py) | Helper that transforms benchmark output into a CSV/plot for visual analysis of routing overhead across scenarios |

## Summary

- Switchyard's routing overhead is typically only a few milliseconds per request, dominated by the classifier's decision cost.
- The framework is fully async, scales linearly with concurrency, and preserves backend error semantics.
- Token throughput is backend-bound since the payload passes through unmodified.
- Memory limits (`MAX_MATERIALIZED_REQUESTS`, `MAX_MATERIALIZED_BYTES`) protect against OOM during load testing.
- The built-in benchmark suite in [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) provides reproducible verification for any new routing algorithm or deployment configuration.

## Frequently Asked Questions

### Does Switchyard add significant latency to LLM requests?

No. Benchmarks show typical routing overhead of just a few milliseconds per request, as measured by `routing_overhead_avg_ms` in the benchmark script. Compared to LLM inference times — often hundreds of milliseconds or more — this overhead is negligible.

### How does Switchyard handle high concurrency?

Switchyard is built on an async architecture and scales linearly with the `concurrency` setting in the benchmark harness. Performance remains stable up to the limits of the underlying HTTP server and event loop.

### Does Switchyard modify the LLM payload?

No. Switchyard forwards the request payload unmodified. As a result, token throughput is dominated by the backend model, and the benchmark verifies that the routing layer adds no extra token-processing bottleneck.

### What do `MAX_MATERIALIZED_REQUESTS` and `MAX_MATERIALIZED_BYTES` do in the benchmark?

These two constants cap the number of in-memory requests and the total byte size of materialized traffic, respectively. They prevent out-of-memory conditions during load testing while still generating realistic traffic patterns.

### How can I reproduce Switchyard's performance numbers?

Install dev dependencies with `uv sync --group dev`, start a `switchyard-server` instance, then run `python -m scripts.benchmark_routing_algorithms` with your base URL, concurrency, and request count. The script emits a full markdown report plus an optional CSV for both routing overhead and throughput analysis.