Switchyard Performance Optimizations: 8 Ways to Speed Up LLM Routing

Switchyard provides eight configurable performance optimizations—including intelligent routing algorithms, token reuse, response caching, and LLM-judge shortcuts—that reduce latency and maximize throughput for LLM request routing.

Switchyard, developed under the NVIDIA-NeMo organization, is an open-source LLM routing engine designed to minimize latency while handling high-throughput inference workloads. Understanding the available performance optimizations allows operators to significantly reduce compute costs and improve response times through strategic configuration and architectural choices.

Routing Algorithm Selection for Optimal Performance

Switchyard implements multiple routing strategies in crates/libsy/src/core/algorithm.rs, where each algorithm implements the route method. Selecting the appropriate strategy is the primary lever for performance tuning, as documented in docs/routing_algorithms/overview.md.

Stage Router (Fastest Path)

The stage_router algorithm provides the fastest path for most workloads by routing requests to cheaper models unless specific hints indicate a need for stronger capabilities. According to docs/routing_algorithms/stage_router_routing.md, this avoids expensive LLM calls when simpler models suffice, cutting compute time significantly.

Configure this optimization in your TOML deployment file:

[routes.smart]
id = "smart"
type = "stage_router"          # optimization: stage-router routing

subagents = false
default_target = "weak"        # route to cheap model unless stage hints indicate otherwise

Advisor Gate (Quality Enforcement)

The advisor_gate algorithm uses a lightweight reviewer LLM to validate "done" claims before returning final answers. This allows the main model to operate at maximum speed while still enforcing output quality through secondary validation.

Token Reuse and Response Caching

Switchyard provides two distinct caching layers that prevent redundant computation: token-level reuse for conversational context and full-response memoization for identical prompts.

Prefill Router Token Reuse

The prefill-router crate enables token reuse across consecutive turns when the same model is called repeatedly. Implemented in crates/prefill-router/build.rs, this optimization reduces the computational overhead of regenerating context by maintaining a reuse_window of recent tokens.

Configure token reuse through the reuse_window parameter:

[routes.chat]
id = "chat"
type = "prefill_router"
target = "strong"
reuse_window = 10             # Re-use the last 10 tokens when possible

Server-Side Response Caching

For identical requests (model ID + prompt), Switchyard memoizes responses at the server level. The cache eligibility logic in crates/switchyard-server/src/stats/cache_eligibility.rs determines which requests qualify for caching based on configuration parameters.

Enable response caching in your server configuration:

[server.cache]
enabled = true                # Turn on memoisation of identical requests

max_size = "2GiB"
ttl_seconds = 300

Early Exit Patterns and Validation Shortcuts

Switchyard implements shortcut mechanisms that terminate expensive operations early when success is impossible or quality is already assured.

LLM Judge Shortcuts

The LLM-judge optimization, located in crates/libsy/src/algorithms/util/llm_judge.rs, enables a judge LLM to reject obviously unsuitable requests before invoking the full production model. This prevents expensive model runs for requests that violate policies or contain malformed inputs.

Advisor-Gate Gating

The advisor-gate mechanism provides a lightweight validation step that checks completion quality without blocking the main inference pipeline. As documented in docs/routing_algorithms/advisor_gate_routing.md, this optimization maintains throughput by offloading quality checks to smaller, faster models.

Streaming and Asynchronous Optimizations

Switchyard minimizes perceived latency through async streaming implemented in crates/switchyard-server/src/sse.rs. The server transmits partial results as Server-Sent Events (SSE), allowing clients to begin processing as soon as the first token arrives. This technique overlaps compute and network latency, significantly improving time-to-first-token metrics.

Observability and Automated Benchmarking

Effective performance optimization requires measurement. Switchyard provides built-in metrics and automated testing harnesses to identify bottlenecks.

Metric Collection

The server collects throughput, latency, and cache-hit ratio statistics through the instrumentation in crates/switchyard-server/src/stats.rs. These metrics enable automatic adjustment of routing weights and identification of performance regressions.

AIPerf Soak Testing

The switchyard-soak crate provides an automated performance-measurement harness. Located in crates/switchyard-soak/README.md, this tool generates realistic traffic patterns and measures per-algorithm latency to produce optimization reports.

Run comprehensive benchmarks using:


# Export a scenario manifest and let the Python runner measure latency

switchyard-soak --export-scenarios scenarios.toml \
                --output-dir perf-report \
                --mode benchmark
python -m aiperf run --manifest perf-report/manifest.json

Configuring Performance Optimizations in TOML

All optimizations are configuration-driven through the TOML deployment file consumed by the switchyard-server binary. Below is a complete example combining multiple optimizations:


# Complete performance-optimized configuration

[switchyard]
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.strong]
id = "openai/gpt-4o"
llm_client = "openrouter"

[targets.weak]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"

# Optimization 1: Stage-router for intelligent model selection

[routes.smart]
id = "smart"
type = "stage_router"
subagents = false
default_target = "weak"

# Optimization 2: Prefill-router for token reuse

[routes.chat]
id = "chat"
type = "prefill_router"
target = "strong"
reuse_window = 10

# Optimization 3: Server configuration with caching

[server]
host = "127.0.0.1"
port = 4000

[server.cache]
enabled = true
max_size = "2GiB"
ttl_seconds = 300

Summary

  • Switchyard provides eight distinct performance optimizations configurable through TOML files and source code modifications.
  • Routing algorithms like stage_router and advisor_gate minimize expensive model calls through intelligent request classification.
  • Token reuse via the prefill-router and response caching eliminate redundant computation for repeated or identical prompts.
  • LLM-judge shortcuts and advisor-gate gating prevent wasted compute on invalid or low-quality requests.
  • Async SSE streaming reduces perceived latency by overlapping network and compute time.
  • Built-in observability and the AIPerf soak-test harness provide the measurement infrastructure necessary for continuous optimization.

Frequently Asked Questions

What is the fastest routing algorithm available in Switchyard?

The stage_router algorithm provides the fastest path for most workloads according to the source code in crates/libsy/src/core/algorithm.rs. It routes requests to cheaper models by default, only escalating to expensive models when specific stage hints indicate a requirement for stronger capabilities. This minimizes compute time and API costs for the majority of requests.

How does the prefill router optimize token usage?

The prefill router reuses tokens from previous conversation turns when calling the same model repeatedly. Implemented in crates/prefill-router/build.rs, this optimization reduces the computational overhead of regenerating context by maintaining a reuse_window of recent tokens. This is particularly effective for chat applications with multi-turn conversations.

Can multiple performance optimizations be used simultaneously?

Yes. Switchyard optimizations are designed to work composably. You can enable stage_router for intelligent routing, prefill_router for token reuse, and server.cache for response memoization within the same deployment configuration. The TOML configuration file allows each optimization to be tuned independently while operating concurrently.

How do I measure the impact of these optimizations on my deployment?

Use the switchyard-soak harness documented in crates/switchyard-soak/README.md to generate realistic traffic patterns and measure per-algorithm latency. Additionally, enable the metrics collection in crates/switchyard-server/src/stats.rs to monitor throughput, latency, and cache-hit ratios in production. These measurements allow you to calculate the specific latency reduction and throughput improvement for your specific workload.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →