Switchyard
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
Learn how Switchyard isolates routing overhead during soak testing. Discover how measuring latency delta reveals Switchyard specific processing time vs direct arm latency.
How the closed_book_proxy Integration in Switchyard Uses an Allowlist Rewriter to Restrict Agent Traffic During BenchmarkingLearn how Switchyard's closed_book_proxy uses an allowlist rewriter with Mitmproxy to block unauthorized agent traffic during benchmarking, ensuring data integrity and efficient testing.
How to Reproduce Terminal-Bench 2.1 Benchmark Results for Switchyard and Configure Escalation DeploymentReproduce Terminal-Bench 2.1 benchmark results for Switchyard. Learn how to prepare the dataset, establish a baseline, and configure escalation deployment using the provided TOML file.
Purpose of fallback_base_url in Switchyard: Handling Unmatched HTTP RequestsLearn the purpose of fallback_base_url in Switchyard for handling unmatched HTTP requests. Discover how this catch-all endpoint proxies requests when no route matches. Optimize your LLM routing.
How a Passthrough Route in Switchyard Registers a Single Target Model ID Without Making a Routing DecisionLearn how Switchyard's passthrough route registers a single target model ID, bypassing routing decisions. Optimize your NLU pipeline by forwarding requests directly.
Python Embedding Path vs Rust Algorithm::run_stream API in Switchyard: Key DifferencesUnderstand the key differences between Python embedding path (Step/LlmResponse) and Rust Algorithm::run_stream API in NVIDIA Switchyard. Learn how Python offers a type-checked facade over the Rust core.
How Switchyard Routes `reasoning_effort` and `extra_body` Parameters to the NetworkDiscover how Switchyard routes reasoning_effort and extra_body parameters to the network. Learn when these parameters are dropped due to capability validation or existing request keys.
Switchyard Runner Duplicate Model ID Warning: Causes and SolutionsResolve Switchyard runner duplicate model ID warnings. Learn why this occurs on the same llm_client and how to fix it for deterministic routing.
Prometheus Counters Emitted by Switchyard-Server: Routing Overhead vs Latency BreakdownExplore Switchyard-server's Prometheus counters. Understand how switchyard.routing_overhead_ms and switchyard.model_call_latency_ms differentiate routing overhead from model latency for optimized performance.
HTTP Endpoints Exposed by the Standalone Switchyard Proxy: Complete API ReferenceExplore the standalone Switchyard proxy's HTTP endpoints including OpenAI compatible chat completions, Anthropic APIs, and Prometheus metrics. Discover each endpoint's function in our API reference.
How Switchyard’s LiteLLM Integration Connects Routing to LiteLLM’s Router/ProxyDiscover how Switchyard's LiteLLM integration connects routing to LiteLLM's Router/proxy. Learn how plugins select optimal models and rewrite requests for seamless AI routing.
How to Build, Package, and Configure the NVIDIA NeMo Relay Plugin for Switchyard to Use a Specific routes.toml DeploymentLearn to build, package, and configure the NVIDIA NeMo Relay plugin for Switchyard. Point your routes.toml deployment and dynamically route LLM requests with this Rust cdylib.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →