How to Configure Multiple LLM Clients with Different Providers in One Switchyard Deployment
Switchyard enables you to define multiple LLM clients from different providers—such as OpenAI, Anthropic, and NVIDIA—in a single TOML configuration file, then route requests between them dynamically using targets and stage routers.
Switchyard is an open-source LLM serving framework developed by NVIDIA-NeMo that unifies disparate AI providers under one OpenAI-compatible endpoint. By leveraging a declarative TOML configuration format, you can configure multiple LLM clients with different providers in one deployment, allowing a single server process to intelligently route traffic across heterogeneous backends based on cost, latency, or capability.
Configuration Architecture Overview
Switchyard’s deployment model relies on three interconnected configuration layers defined in one TOML file. Understanding how llm_clients, targets, and routes interact is essential for building multi-provider setups.
LLM Client Definitions
The top-level table [llm_clients.<name>] declares how Switchyard communicates with a specific provider. Each client requires three fields:
format: Selects the request/response schema (e.g.,openai_chat,anthropic_messages)base_url: The provider’s API endpointapi_key_env: The name of the environment variable containing the secret key
According to the reference configuration in [dev-server/config.toml](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/dev-server/config.toml), these definitions abstract provider-specific protocols into a unified interface that the server can consume.
Target Definitions
Targets bind specific model IDs to LLM clients using the [targets.<name>] table. The critical field is llm_client, which must match a name defined in the llm_clients section above. This creates a bridge between a logical model identifier (like claude-3-opus) and the physical connection details required to reach Anthropic’s servers.
Route Compositions
Routes determine which target handles an incoming request. The [routes.<name>] tables can reference targets that use different llm_clients, enabling sophisticated hybrid routing strategies. For example, a stage router can attempt a cost-effective NVIDIA model first, then escalate to a more capable Anthropic model if confidence scores fall below a threshold.
Complete Multi-Provider Configuration Example
Below is a production-ready TOML configuration that simultaneously connects to OpenAI, Anthropic, and NVIDIA inference endpoints within one deployment.
schema_version = 1
# ---- LLM client definitions -------------------------------------------------
[llm_clients.openai]
format = "openai_chat"
base_url = "https://api.openai.com/v1"
api_key_env = "OPENAI_API_KEY"
[llm_clients.anthropic]
format = "anthropic_messages"
base_url = "https://api.anthropic.com"
api_key_env = "ANTHROPIC_API_KEY"
[llm_clients.nvidia_hub]
format = "openai_chat"
base_url = "https://inference-api.nvidia.com/v1"
api_key_env = "NVIDIA_API_KEY"
# ---- Model targets ---------------------------------------------------------
[targets.gpt4]
id = "openai/gpt-4"
llm_client = "openai"
[targets.claude]
id = "anthropic/claude-3-opus-20240229"
llm_client = "anthropic"
[targets.nemo_llama]
id = "nvidia/lama-13b"
llm_client = "nvidia_hub"
# ---- Routes ---------------------------------------------------------------
[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "claude"
efficient_target = "nemo_llama"
picker = "efficient_first"
confidence_threshold = 0.5
[routes.random]
id = "switchyard/random"
type = "random"
targets = ["gpt4", "claude", "nemo_llama"]
Lines 5–13 define three distinct providers using different API formats. Lines 16–24 bind each model to its respective client, while lines 27–38 create routes that dynamically select between these heterogeneous backends.
Launching the Multi-Provider Server
After defining the configuration, export the API keys referenced in your TOML file and launch the server:
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-ant..."
export NVIDIA_API_KEY="sk-nvidia..."
switchyard launch claude --config my_deployment.toml
The single switchyard process exposes a unified OpenAI-compatible endpoint at http://localhost:4000/v1. When a request matches the stage route, the server first attempts the cost-effective nvidia_hub target, then automatically fails over to the anthropic target if the confidence threshold is not met.
Internal Implementation Details
When the server starts, the binary at switchyard-server parses the TOML configuration in crates/switchyard-server/src/config.rs. At line 57, it constructs a llm_clients: BTreeMap<String, LlmClientConfig> that maps client names to their connection parameters.
Upon receiving a request, the server resolves the target specified in the route. At line 129 of the same file, it looks up the target’s llm_client field and retrieves the corresponding LlmClientConfig from the map. This configuration drives the HTTP request to the provider, serializing the payload according to the specified format and injecting the API key from the designated environment variable.
Client Integration
Applications consume the multi-provider deployment through Switchyard’s unified API without knowing which backend serves the request:
import switchyard
client = switchyard.Client(
base_url="http://localhost:4000/v1",
api_key="any-key" # Switchyard endpoint validates internally
)
response = client.chat_completions.create(
model="switchyard/stage", # Uses the multi-provider stage router
messages=[{"role": "user", "content": "Explain quantum entanglement."}]
)
print(response.choices[0].message.content)
The client code remains agnostic to provider-specific implementations while Switchyard handles the protocol translation and failover logic defined in the TOML configuration.
Summary
- Declarative Configuration: Define unlimited LLM providers using
[llm_clients.<name>]tables withformat,base_url, andapi_key_envfields. - Model Binding: Associate specific models with clients via
[targets.<name>]using thellm_clientreference field. - Hybrid Routing: Compose
[routes.<name>]that mix targets from different providers, enabling cost-effective or capability-based failover strategies. - Unified Runtime: A single
switchyard launchprocess serves all configured providers through one OpenAI-compatible endpoint.
Frequently Asked Questions
How many different LLM providers can I configure in a single deployment?
You can define an unlimited number of providers in one TOML file, restricted only by system memory and environment variable availability. The BTreeMap<String, LlmClientConfig> structure in crates/switchyard-server/src/config.rs holds all client definitions in memory at runtime.
Is it possible to use the same model ID from different providers?
Yes, by defining multiple targets with different llm_client values but distinct target names. For example, you could have targets.gpt4_openai and targets.gpt4_azure both pointing to GPT-4 models but using different client configurations and base URLs.
What happens if one provider's API key environment variable is not set?
The Switchyard server will start successfully, but any requests routed to targets using that client will fail at runtime with an authentication error. The server validates connectivity only when attempting to forward requests, not during the initial configuration load.
Can I route a single request to multiple providers simultaneously?
The built-in stage router and random router select one target per request. To query multiple providers simultaneously, you would need to implement client-side parallelism or develop a custom router plugin that dispatches to multiple targets and aggregates responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →