# How to Configure Multiple LLM Clients with Different Providers in One Switchyard Deployment

> Learn to configure multiple LLM clients from providers like OpenAI Anthropic and NVIDIA in one Switchyard deployment Route requests dynamically using targets and stage routers

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Switchyard enables you to define multiple LLM clients from different providers—such as OpenAI, Anthropic, and NVIDIA—in a single TOML configuration file, then route requests between them dynamically using targets and stage routers.**

Switchyard is an open-source LLM serving framework developed by NVIDIA-NeMo that unifies disparate AI providers under one OpenAI-compatible endpoint. By leveraging a declarative TOML configuration format, you can configure multiple LLM clients with different providers in one deployment, allowing a single server process to intelligently route traffic across heterogeneous backends based on cost, latency, or capability.

## Configuration Architecture Overview

Switchyard’s deployment model relies on three interconnected configuration layers defined in one TOML file. Understanding how `llm_clients`, `targets`, and `routes` interact is essential for building multi-provider setups.

### LLM Client Definitions

The top-level table **`[llm_clients.<name>]`** declares how Switchyard communicates with a specific provider. Each client requires three fields:

- **`format`**: Selects the request/response schema (e.g., `openai_chat`, `anthropic_messages`)
- **`base_url`**: The provider’s API endpoint
- **`api_key_env`**: The name of the environment variable containing the secret key

According to the reference configuration in [[`dev-server/config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/dev-server/config.toml)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/dev-server/config.toml), these definitions abstract provider-specific protocols into a unified interface that the server can consume.

### Target Definitions

Targets bind specific model IDs to LLM clients using the **`[targets.<name>]`** table. The critical field is **`llm_client`**, which must match a name defined in the `llm_clients` section above. This creates a bridge between a logical model identifier (like `claude-3-opus`) and the physical connection details required to reach Anthropic’s servers.

### Route Compositions

Routes determine which target handles an incoming request. The **`[routes.<name>]`** tables can reference targets that use different `llm_clients`, enabling sophisticated hybrid routing strategies. For example, a stage router can attempt a cost-effective NVIDIA model first, then escalate to a more capable Anthropic model if confidence scores fall below a threshold.

## Complete Multi-Provider Configuration Example

Below is a production-ready TOML configuration that simultaneously connects to OpenAI, Anthropic, and NVIDIA inference endpoints within one deployment.

```toml
schema_version = 1

# ---- LLM client definitions -------------------------------------------------

[llm_clients.openai]
format = "openai_chat"
base_url = "https://api.openai.com/v1"
api_key_env = "OPENAI_API_KEY"

[llm_clients.anthropic]
format = "anthropic_messages"
base_url = "https://api.anthropic.com"
api_key_env = "ANTHROPIC_API_KEY"

[llm_clients.nvidia_hub]
format = "openai_chat"
base_url = "https://inference-api.nvidia.com/v1"
api_key_env = "NVIDIA_API_KEY"

# ---- Model targets ---------------------------------------------------------

[targets.gpt4]
id = "openai/gpt-4"
llm_client = "openai"

[targets.claude]
id = "anthropic/claude-3-opus-20240229"
llm_client = "anthropic"

[targets.nemo_llama]
id = "nvidia/lama-13b"
llm_client = "nvidia_hub"

# ---- Routes ---------------------------------------------------------------

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "claude"
efficient_target = "nemo_llama"
picker = "efficient_first"
confidence_threshold = 0.5

[routes.random]
id = "switchyard/random"
type = "random"
targets = ["gpt4", "claude", "nemo_llama"]

```

Lines 5–13 define three distinct providers using different API formats. Lines 16–24 bind each model to its respective client, while lines 27–38 create routes that dynamically select between these heterogeneous backends.

## Launching the Multi-Provider Server

After defining the configuration, export the API keys referenced in your TOML file and launch the server:

```bash
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-ant..."
export NVIDIA_API_KEY="sk-nvidia..."

switchyard launch claude --config my_deployment.toml

```

The single `switchyard` process exposes a unified OpenAI-compatible endpoint at `http://localhost:4000/v1`. When a request matches the `stage` route, the server first attempts the cost-effective `nvidia_hub` target, then automatically fails over to the `anthropic` target if the confidence threshold is not met.

## Internal Implementation Details

When the server starts, the binary at `switchyard-server` parses the TOML configuration in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs). At line 57, it constructs a **`llm_clients: BTreeMap<String, LlmClientConfig>`** that maps client names to their connection parameters.

Upon receiving a request, the server resolves the target specified in the route. At line 129 of the same file, it looks up the target’s `llm_client` field and retrieves the corresponding `LlmClientConfig` from the map. This configuration drives the HTTP request to the provider, serializing the payload according to the specified `format` and injecting the API key from the designated environment variable.

## Client Integration

Applications consume the multi-provider deployment through Switchyard’s unified API without knowing which backend serves the request:

```python
import switchyard

client = switchyard.Client(
    base_url="http://localhost:4000/v1",
    api_key="any-key"  # Switchyard endpoint validates internally

)

response = client.chat_completions.create(
    model="switchyard/stage",  # Uses the multi-provider stage router

    messages=[{"role": "user", "content": "Explain quantum entanglement."}]
)

print(response.choices[0].message.content)

```

The client code remains agnostic to provider-specific implementations while Switchyard handles the protocol translation and failover logic defined in the TOML configuration.

## Summary

- **Declarative Configuration**: Define unlimited LLM providers using `[llm_clients.<name>]` tables with `format`, `base_url`, and `api_key_env` fields.
- **Model Binding**: Associate specific models with clients via `[targets.<name>]` using the `llm_client` reference field.
- **Hybrid Routing**: Compose `[routes.<name>]` that mix targets from different providers, enabling cost-effective or capability-based failover strategies.
- **Unified Runtime**: A single `switchyard launch` process serves all configured providers through one OpenAI-compatible endpoint.

## Frequently Asked Questions

### How many different LLM providers can I configure in a single deployment?

You can define an unlimited number of providers in one TOML file, restricted only by system memory and environment variable availability. The `BTreeMap<String, LlmClientConfig>` structure in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs) holds all client definitions in memory at runtime.

### Is it possible to use the same model ID from different providers?

Yes, by defining multiple targets with different `llm_client` values but distinct target names. For example, you could have `targets.gpt4_openai` and `targets.gpt4_azure` both pointing to GPT-4 models but using different client configurations and base URLs.

### What happens if one provider's API key environment variable is not set?

The Switchyard server will start successfully, but any requests routed to targets using that client will fail at runtime with an authentication error. The server validates connectivity only when attempting to forward requests, not during the initial configuration load.

### Can I route a single request to multiple providers simultaneously?

The built-in stage router and random router select one target per request. To query multiple providers simultaneously, you would need to implement client-side parallelism or develop a custom router plugin that dispatches to multiple targets and aggregates responses.