How Switchyard Routes LLM Requests Across Different Providers
Switchyard routes LLM requests across different providers using a hybrid Python-Rust architecture where configurable algorithms in the libsy crate select targets based on model names, tasks, or custom strategies, then delegate to provider-specific HTTP clients.
Switchyard, NVIDIA's open-source LLM serving platform, abstracts away provider differences through a sophisticated routing layer. This system inspects incoming requests and directs them to the appropriate backend—whether OpenAI, Anthropic, NVIDIA, or OpenRouter—using algorithms defined in the libsy crate and exposed via switchyard_rust bindings.
The Hybrid Architecture: Python Frontend to Rust Core
The routing system spans two languages for performance and ergonomics. Python wrappers in switchyard/libsy/algorithms.py provide the user-facing API, while the Rust libsy crate handles the actual routing logic. The switchyard_rust bindings bridge these layers, allowing Python code to invoke high-performance Rust algorithms without overhead.
This architecture separates configuration (handled in Python) from execution (optimized in Rust). When you instantiate a SwitchyardClient with a specific router, you're selecting a Rust-implemented strategy that operates on a RoutingTable built from your TOML deployment configuration.
Routing Algorithms in switchyard/libsy/algorithms.py
The decision logic resides in switchyard/libsy/algorithms.py, which exposes four primary strategies as thin Python wrappers around Rust implementations. Each algorithm implements a different selection criteria for mapping requests to provider targets.
Model-Based Classification with llm_classifier
The llm_classifier algorithm routes requests based on the model identifier string (e.g., openai/gpt-4, anthropic/claude-3). This is the default strategy, parsing the model name to determine the appropriate provider backend. It performs a lookup in the RoutingTable to match the namespace prefix against configured targets.
Task-Based Routing with llm_task_classifier
For deployments where the same model name might handle different operations, llm_task_classifier inspects the request type—chat, completion, embedding, or moderation—to select the optimal target. This allows fine-grained control over which provider handles specific inference tasks.
Random and Staged failover Strategies
The random algorithm selects targets probabilistically, useful for A/B testing or load distribution across equivalent providers. For complex deployments, stage_router composes multiple strategies into a pipeline, attempting a primary router first and falling back to secondary options if the primary returns no valid target.
Provider Configuration and the Routing Table
Targets are defined in a TOML deployment file that the native server reads at startup. Each entry specifies the provider type, model identifier, credentials, and endpoint URLs. The server implementation in crates/switchyard-server/src/deployment.rs (implicitly handling the TOML parsing) constructs a RoutingTable from these definitions.
The RoutingTable serves as the lookup source for all routing algorithms. When Rust code in switchyard_rust/libsy.py invokes a routing decision, it passes this immutable table to the selected algorithm, which returns a target ID corresponding to the chosen provider configuration.
The Request Journey: From Client to Provider
Understanding the exact flow reveals how Switchyard achieves seamless provider abstraction:
- The Python frontend receives a request via
client.chat()or similar methods - The request passes to
switchyard_rust/libsy.py, which calls into the Rustlibsyrouter - The router invokes the selected algorithm (e.g.,
llm_classifier) against theRoutingTable - The algorithm returns the target ID for the appropriate provider
crates/libsy-llm-client/src/lib.rstranslates the generic Switchyard request into the provider-specific HTTP payload format- The HTTP client dispatches the request to the actual backend (OpenAI, Anthropic, etc.)
This pipeline ensures that provider-specific quirks—such as differing JSON schemas or authentication headers—are handled in the translation layer rather than forcing users to manage multiple client libraries.
Observability and Routing Metrics
After dispatch, crates/switchyard-server/src/routing_log.rs records detailed telemetry about each routing decision. The system tracks routing overhead, model call counts, and error rates. These metrics can be exported as JSON via the --routing-stats-json flag and visualized using built-in reporting scripts.
This logging occurs in the Rust layer for minimal performance impact, capturing the exact target selected and the time spent in routing logic versus actual model inference.
Code Examples for Custom Routing Strategies
Default Model-Based Routing
The simplest usage relies on the llm_classifier to parse model names and route accordingly:
import asyncio
from switchyard import SwitchyardClient
async def main():
client = SwitchyardClient()
response = await client.chat(
model="openai/gpt-4",
messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
)
print(response["choices"][0]["message"]["content"])
asyncio.run(main())
Random Selection for A/B Testing
To distribute load randomly across configured targets:
from switchyard.libsy import algorithms as routing
from switchyard import SwitchyardClient
# Create a random router
router = routing.random()
# Attach the router to the client
client = SwitchyardClient(router=router)
# The request will be sent to a randomly chosen target defined in the TOML config
await client.chat(model="any-model", messages=[...])
Composed Routing with Fallback
Use stage_router to attempt model-based routing first, then fall back to random selection if no specific match exists:
from switchyard.libsy import algorithms as routing
stage = routing.stage_router(
primary=routing.llm_classifier(),
fallback=routing.random(),
)
client = SwitchyardClient(router=stage)
await client.chat(model="openai/gpt-4", messages=[...])
Summary
- Switchyard implements routing through a Rust core (
libsycrate) exposed via Python bindings inswitchyard_rust/libsy.pyfor high performance with Pythonic ergonomics. - Four primary algorithms handle selection:
llm_classifier(by model name),llm_task_classifier(by task type),random(probabilistic), andstage_router(composable pipelines). - Provider targets are defined in TOML deployment files and loaded into an immutable
RoutingTableparsed by the server implementation. - The HTTP translation layer in
crates/libsy-llm-client/src/lib.rsconverts generic requests to provider-specific formats after routing decisions are made. - Routing decisions and performance metrics are logged via
crates/switchyard-server/src/routing_log.rs, supporting JSON export and visualization.
Frequently Asked Questions
How does Switchyard decide which LLM provider to use for a request?
Switchyard applies a configurable routing algorithm defined in switchyard/libsy/algorithms.py that inspects the request model name, task type, or custom criteria. The algorithm looks up the matching target in a RoutingTable built from the TOML deployment configuration and returns the specific provider ID to handle the request.
Can I implement custom routing logic beyond the built-in algorithms?
Yes. While the built-in algorithms in switchyard/libsy/algorithms.py cover most use cases, you can compose custom logic using the stage_router to chain multiple strategies, or extend the Python wrapper interface to interact with the underlying Rust libsy routing engine for entirely bespoke selection logic.
Where are routing decisions logged in the Switchyard codebase?
Routing decisions are recorded in crates/switchyard-server/src/routing_log.rs, which captures the selected target, routing overhead latency, and error states. These metrics can be exported as JSON using the --routing-stats-json command-line flag for external analysis and monitoring.
What file format configures provider targets in Switchyard?
Provider targets are configured using TOML deployment files that specify each target's provider type (OpenAI, Anthropic, etc.), model identifiers, and authentication credentials. The server parses these files at startup to construct the RoutingTable used by all routing algorithms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →