# Switchyard Performance Optimizations: 8 Ways to Speed Up LLM Routing

> Discover 8 Switchyard performance optimizations to speed up LLM routing. Reduce latency and maximize throughput with intelligent routing, token reuse, and more.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: performance
- Published: 2026-09-11

---

**Switchyard provides eight configurable performance optimizations—including intelligent routing algorithms, token reuse, response caching, and LLM-judge shortcuts—that reduce latency and maximize throughput for LLM request routing.**

Switchyard, developed under the NVIDIA-NeMo organization, is an open-source LLM routing engine designed to minimize latency while handling high-throughput inference workloads. Understanding the available performance optimizations allows operators to significantly reduce compute costs and improve response times through strategic configuration and architectural choices.

## Routing Algorithm Selection for Optimal Performance

Switchyard implements multiple routing strategies in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), where each algorithm implements the `route` method. Selecting the appropriate strategy is the primary lever for performance tuning, as documented in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md).

### Stage Router (Fastest Path)

The **stage_router** algorithm provides the fastest path for most workloads by routing requests to cheaper models unless specific hints indicate a need for stronger capabilities. According to [`docs/routing_algorithms/stage_router_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/stage_router_routing.md), this avoids expensive LLM calls when simpler models suffice, cutting compute time significantly.

Configure this optimization in your TOML deployment file:

```toml
[routes.smart]
id = "smart"
type = "stage_router"          # optimization: stage-router routing

subagents = false
default_target = "weak"        # route to cheap model unless stage hints indicate otherwise

```

### Advisor Gate (Quality Enforcement)

The **advisor_gate** algorithm uses a lightweight reviewer LLM to validate "done" claims before returning final answers. This allows the main model to operate at maximum speed while still enforcing output quality through secondary validation.

## Token Reuse and Response Caching

Switchyard provides two distinct caching layers that prevent redundant computation: token-level reuse for conversational context and full-response memoization for identical prompts.

### Prefill Router Token Reuse

The **prefill-router** crate enables token reuse across consecutive turns when the same model is called repeatedly. Implemented in [`crates/prefill-router/build.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/prefill-router/build.rs), this optimization reduces the computational overhead of regenerating context by maintaining a `reuse_window` of recent tokens.

Configure token reuse through the `reuse_window` parameter:

```toml
[routes.chat]
id = "chat"
type = "prefill_router"
target = "strong"
reuse_window = 10             # Re-use the last 10 tokens when possible

```

### Server-Side Response Caching

For identical requests (model ID + prompt), Switchyard memoizes responses at the server level. The cache eligibility logic in [`crates/switchyard-server/src/stats/cache_eligibility.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats/cache_eligibility.rs) determines which requests qualify for caching based on configuration parameters.

Enable response caching in your server configuration:

```toml
[server.cache]
enabled = true                # Turn on memoisation of identical requests

max_size = "2GiB"
ttl_seconds = 300

```

## Early Exit Patterns and Validation Shortcuts

Switchyard implements shortcut mechanisms that terminate expensive operations early when success is impossible or quality is already assured.

### LLM Judge Shortcuts

The **LLM-judge** optimization, located in [`crates/libsy/src/algorithms/util/llm_judge.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/util/llm_judge.rs), enables a judge LLM to reject obviously unsuitable requests before invoking the full production model. This prevents expensive model runs for requests that violate policies or contain malformed inputs.

### Advisor-Gate Gating

The **advisor-gate** mechanism provides a lightweight validation step that checks completion quality without blocking the main inference pipeline. As documented in [`docs/routing_algorithms/advisor_gate_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/advisor_gate_routing.md), this optimization maintains throughput by offloading quality checks to smaller, faster models.

## Streaming and Asynchronous Optimizations

Switchyard minimizes perceived latency through **async streaming** implemented in [`crates/switchyard-server/src/sse.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/sse.rs). The server transmits partial results as **Server-Sent Events (SSE)**, allowing clients to begin processing as soon as the first token arrives. This technique overlaps compute and network latency, significantly improving time-to-first-token metrics.

## Observability and Automated Benchmarking

Effective performance optimization requires measurement. Switchyard provides built-in metrics and automated testing harnesses to identify bottlenecks.

### Metric Collection

The server collects throughput, latency, and cache-hit ratio statistics through the instrumentation in [`crates/switchyard-server/src/stats.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats.rs). These metrics enable automatic adjustment of routing weights and identification of performance regressions.

### AIPerf Soak Testing

The **switchyard-soak** crate provides an automated performance-measurement harness. Located in [`crates/switchyard-soak/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/README.md), this tool generates realistic traffic patterns and measures per-algorithm latency to produce optimization reports.

Run comprehensive benchmarks using:

```bash

# Export a scenario manifest and let the Python runner measure latency

switchyard-soak --export-scenarios scenarios.toml \
                --output-dir perf-report \
                --mode benchmark
python -m aiperf run --manifest perf-report/manifest.json

```

## Configuring Performance Optimizations in TOML

All optimizations are configuration-driven through the TOML deployment file consumed by the `switchyard-server` binary. Below is a complete example combining multiple optimizations:

```toml

# Complete performance-optimized configuration

[switchyard]
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

# Optimization 1: Stage-router for intelligent model selection

[routes.smart]
id = "smart"
type = "stage_router"
subagents = false
default_target = "weak"

# Optimization 2: Prefill-router for token reuse

[routes.chat]
id = "chat"
type = "prefill_router"
target = "strong"
reuse_window = 10

# Optimization 3: Server configuration with caching

[server]
host = "127.0.0.1"
port = 4000

[server.cache]
enabled = true
max_size = "2GiB"
ttl_seconds = 300

```

## Summary

- **Switchyard** provides eight distinct performance optimizations configurable through TOML files and source code modifications.
- **Routing algorithms** like `stage_router` and `advisor_gate` minimize expensive model calls through intelligent request classification.
- **Token reuse** via the prefill-router and **response caching** eliminate redundant computation for repeated or identical prompts.
- **LLM-judge shortcuts** and **advisor-gate gating** prevent wasted compute on invalid or low-quality requests.
- **Async SSE streaming** reduces perceived latency by overlapping network and compute time.
- **Built-in observability** and the **AIPerf soak-test harness** provide the measurement infrastructure necessary for continuous optimization.

## Frequently Asked Questions

### What is the fastest routing algorithm available in Switchyard?

The **stage_router** algorithm provides the fastest path for most workloads according to the source code in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs). It routes requests to cheaper models by default, only escalating to expensive models when specific stage hints indicate a requirement for stronger capabilities. This minimizes compute time and API costs for the majority of requests.

### How does the prefill router optimize token usage?

The **prefill router** reuses tokens from previous conversation turns when calling the same model repeatedly. Implemented in [`crates/prefill-router/build.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/prefill-router/build.rs), this optimization reduces the computational overhead of regenerating context by maintaining a `reuse_window` of recent tokens. This is particularly effective for chat applications with multi-turn conversations.

### Can multiple performance optimizations be used simultaneously?

Yes. Switchyard optimizations are designed to work composably. You can enable **stage_router** for intelligent routing, **prefill_router** for token reuse, and **server.cache** for response memoization within the same deployment configuration. The TOML configuration file allows each optimization to be tuned independently while operating concurrently.

### How do I measure the impact of these optimizations on my deployment?

Use the **switchyard-soak** harness documented in [`crates/switchyard-soak/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-soak/README.md) to generate realistic traffic patterns and measure per-algorithm latency. Additionally, enable the metrics collection in [`crates/switchyard-server/src/stats.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/stats.rs) to monitor throughput, latency, and cache-hit ratios in production. These measurements allow you to calculate the specific latency reduction and throughput improvement for your specific workload.