Switchyard Performance Considerations: Latency, Throughput, and Scaling Explained
TLDR: Switchyard adds only a few milliseconds of routing overhead per LLM request, scales linearly with concurrency, preserves backend error semantics, and ships with a built-in benchmark suite in scripts/benchmark_routing_algorithms.py to measure every performance dimension.
Switchyard, an open-source orchestration layer from NVIDIA-NeMo/Switchyard, sits between client applications and large-language-model backends, classifying and forwarding each request. While such a routing layer can introduce overhead, the project's design and benchmark tooling are built to keep that impact minimal. This article breaks down the key performance considerations for Switchyard, the metrics that matter, and exactly how to measure them using the repository's own benchmark suite.
The Performance Dimensions Switchyard Tracks
Switchyard's benchmark suite categorizes performance impact across five measurable dimensions. Each maps to a concrete metric collected during load testing:
| Dimension | What it measures | Typical metrics |
|---|---|---|
| Routing latency | Time spent classifying and forwarding each request | routing_overhead_avg_ms |
| Throughput | Requests or tokens processed per second | oha_requests_per_second, aiperf_requests_per_second, aiperf_output_tokens_per_second |
| Error handling | Ability to keep error rates low while routing | aiperf_error_rate, expected_error_rate_min/max |
| Concurrency scaling | Behavior under high parallel load | concurrency, request_rate |
| Resource usage | Memory and request-size limits for materialized traffic | MAX_MATERIALIZED_REQUESTS, MAX_MATERIALIZED_BYTES |
The core benchmark code that collects these metrics lives in scripts/benchmark_routing_algorithms.py, and the accompanying test suite (tests/test_routing_performance_report.py) validates that the reporting logic computes and aggregates each value correctly.
The Five Main Ways Switchyard Affects Performance
1. Routing Overhead
Every request passes through a classifier before reaching the chosen backend. That classification cost is captured in the routing_overhead_avg_ms field, which is defined by the RoutingDelta dataclass in the benchmark script. In practice, benchmarks show overheads of a few milliseconds per request — negligible compared to the inference latency of most LLM backends.
2. Concurrent Request Handling
Switchyard is fully asynchronous and sustains high concurrency. The concurrency parameter in the benchmark harness drives both oha and AIPerf to verify that Switchyard scales linearly, up to the limits of the underlying HTTP server and the Python event loop.
3. Token Throughput
Because Switchyard forwards the payload unmodified, token throughput (aiperf_output_tokens_per_second) is dominated by the backend model, not the router. The benchmark measures the delta between a direct backend call and a routed call to confirm the routing layer adds no extra token-processing bottleneck.
4. Error Propagation
The routing logic must preserve backend error semantics exactly. The test suite compares error counts and rates from a direct backend (compare_to_direct_backend) against those observed through Switchyard, ensuring error amplification stays within a tight bound, as enforced by the expected_error_rate_min/max fields.
5. Request Size Handling
The MAX_MATERIALIZED_REQUESTS constant caps the number of in-memory requests, and MAX_MATERIALIZED_BYTES caps the total byte size. These two limits prevent out-of-memory conditions while still supplying realistic traffic patterns for benchmarks.
Running the Switchyard Benchmark
The benchmark suite ships with the repository in the scripts/ directory. Install development dependencies and launch the server, then run the benchmark:
# Install development dependencies
uv sync --group dev
# Launch a Switchyard server (example config)
switchyard-server --config routes.toml --port 4000
# Run the benchmark (replace <BASE_URL> with your server address)
python -m scripts.benchmark_routing_algorithms \
--base-url http://127.0.0.1:4000 \
--concurrency 100 \
--request-count 1000 \
--output-dir ./benchmark-results
The script then:
- Exports scenarios from the native Rust "soak" tool.
- Executes a short baseline with
ohaand a full load test with AIPerf. - Computes routing overhead, throughput, and latency deltas.
- Emits a markdown report (via the
write_reportfunction) and an optional CSV for downstream analysis.
Interpreting the Benchmark Report
A typical report row from the generated markdown looks like this:
| algorithm | model | scenario | oha_rps | oha_p50 ms | aiperf_rps | aiperf_p50 ms | routing_overhead ms |
|---|---|---|---|---|---|---|---|
random |
switchyard/random |
short-interactive |
120.5 | 0.012 | 115.3 | 0.018 | 2.4 |
Key columns to focus on:
routing_overhead_ms– The extra latency introduced by Switchyard's classifier. Values in the low single-digit milliseconds are typical and expected.oha_rpsvsaiperf_rps– Baseline versus routed requests per-second. A small drop indicates acceptable overhead; a large gap suggests a tuning opportunity.- Error columns –
aiperf_error_rateshould match the direct backend's rate, and the test suite enforces this equality.
Key Source Files for Performance
| File | Role |
|---|---|
scripts/benchmark_routing_algorithms.py |
Main benchmark driver that materializes traffic, runs oha and AIPerf, and writes performance reports |
tests/test_routing_performance_report.py |
Unit test validating the comparison of routed versus direct backend results and ensuring correct aggregation in the report |
scripts/routing_overhead_plot.py |
Helper that transforms benchmark output into a CSV/plot for visual analysis of routing overhead across scenarios |
Summary
- Switchyard's routing overhead is typically only a few milliseconds per request, dominated by the classifier's decision cost.
- The framework is fully async, scales linearly with concurrency, and preserves backend error semantics.
- Token throughput is backend-bound since the payload passes through unmodified.
- Memory limits (
MAX_MATERIALIZED_REQUESTS,MAX_MATERIALIZED_BYTES) protect against OOM during load testing. - The built-in benchmark suite in
scripts/benchmark_routing_algorithms.pyprovides reproducible verification for any new routing algorithm or deployment configuration.
Frequently Asked Questions
Does Switchyard add significant latency to LLM requests?
No. Benchmarks show typical routing overhead of just a few milliseconds per request, as measured by routing_overhead_avg_ms in the benchmark script. Compared to LLM inference times — often hundreds of milliseconds or more — this overhead is negligible.
How does Switchyard handle high concurrency?
Switchyard is built on an async architecture and scales linearly with the concurrency setting in the benchmark harness. Performance remains stable up to the limits of the underlying HTTP server and event loop.
Does Switchyard modify the LLM payload?
No. Switchyard forwards the request payload unmodified. As a result, token throughput is dominated by the backend model, and the benchmark verifies that the routing layer adds no extra token-processing bottleneck.
What do MAX_MATERIALIZED_REQUESTS and MAX_MATERIALIZED_BYTES do in the benchmark?
These two constants cap the number of in-memory requests and the total byte size of materialized traffic, respectively. They prevent out-of-memory conditions during load testing while still generating realistic traffic patterns.
How can I reproduce Switchyard's performance numbers?
Install dev dependencies with uv sync --group dev, start a switchyard-server instance, then run python -m scripts.benchmark_routing_algorithms with your base URL, concurrency, and request count. The script emits a full markdown report plus an optional CSV for both routing overhead and throughput analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →