# How to Configure an LLM Client in Switchyard for OpenRouter

> Learn to configure an LLM client in Switchyard for OpenRouter. Route LLM requests via direct Python or TOML server setup for seamless integration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Switchyard routes LLM requests through a configurable client abstraction that translates internal `libsy` formats into OpenRouter's OpenAI-compatible API, either via direct Python instantiation or TOML-based server configuration.**

Switchyard is a high-performance LLM routing engine developed by NVIDIA that separates request routing logic from provider-specific implementations. When you **configure an LLM client in Switchyard** for **OpenRouter**, you leverage the `LiteLLMSyClient` class to handle API translation and credential management. This client conforms to Switchyard's async `call` contract, enabling seamless integration with the Rust-based routing core while maintaining Python-side flexibility for provider-specific logic.

## Architecture of the OpenRouter Client Integration

Switchyard's architecture decouples the HTTP server from the LLM provider through a **client registry** populated at startup. When routing to OpenRouter, the system chains through three layers: the Switchyard Server receives the request, the `libsy` routing engine selects the appropriate client, and the `LiteLLMSyClient` forwards the normalized request to a LiteLLM gateway configured for OpenRouter's endpoint.

### The LiteLLMSyClient Implementation

The concrete implementation resides in [`examples/experimental/litellm/src/switchyard_litellm/client.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/experimental/litellm/src/switchyard_litellm/client.py). The `LiteLLMSyClient` class (defined around line 319) implements the required interface for Switchyard's routing engine:

```python
class LiteLLMSyClient:
    async def call(self, request: Mapping[str, Any]) -> LlmResponse:
        ...

```

This method converts Switchyard's normalized request format—containing fields like `_messages`, `_tools`, and `_tool_choice`—into a LiteLLM-compatible payload. The conversion helpers (lines ~17-95 in the same file) handle the schema mapping, while the client manages OpenRouter-specific headers and error handling for conditions like `ContextWindowExceededError`.

### Configuration-Driven Client Registration

For server deployments, Switchyard supports declarative configuration via TOML files located in `benchmark/server-configs/`. The `[llm_clients.openrouter]` section (exemplified in [`tb-lite-single-opus-4-7.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tb-lite-single-opus-4-7.toml), lines ~9-13) defines the endpoint and authentication:

```toml
[llm_clients.openrouter]
base_url = "http://localhost:8000"
api_key_env = "OPENROUTER_API_KEY"

```

At server startup, Switchyard reads the `api_key_env` value and injects the corresponding environment variable into the client instance, eliminating the need to hardcode credentials in configuration files.

## Method 1: Direct Client Instantiation in Python

When operating Switchyard as a library or testing integrations locally, instantiate `LiteLLMSyClient` directly. Ensure the `OPENROUTER_API_KEY` environment variable is set, as the client reads this automatically during initialization.

```python
import os
from switchyard_litellm.client import LiteLLMSyClient

# Configure authentication

os.environ["OPENROUTER_API_KEY"] = "sk-your-openrouter-key"

# Initialize client pointing to your LiteLLM gateway

client = LiteLLMSyClient(base_url="http://localhost:8000")

# Construct a libsy-compatible request

request = {
    "model": "openrouter/moonshotai/kimi-k3",
    "messages": [
        {"role": "user", "content": [{"type": "text", "text": "Write a haiku"}]}
    ],
    "max_tokens": 64,
}

# Execute the async call conforming to Switchyard's contract

response = await client.call(request)

print("Model:", response.model)
print("Content:", response.choices[0].message.content)

```

The `call` method returns an `LlmResponse` object containing the model output, token usage, and stop reasons, formatted consistently regardless of the underlying provider.

## Method 2: Server-Side TOML Configuration

For production deployments using the Switchyard server binary, define the client entirely through configuration. Create a TOML file that specifies the OpenRouter client parameters:

```toml

# my_openrouter_config.toml

[llm_clients.openrouter]
base_url = "http://localhost:8000"   # LiteLLM proxy/gateway URL

api_key_env = "OPENROUTER_API_KEY"   # Environment variable name

```

Launch the server with this configuration and the API key exposed in your environment:

```bash
OPENROUTER_API_KEY=sk-your-key \
switchyard-server --config my_openrouter_config.toml --port 4000

```

The server instantiates the `LiteLLMSyClient` automatically, injecting the `base_url` and resolved API key. All HTTP requests hitting the server are then routed through this configured OpenRouter client.

## Routing Requests to the OpenRouter Client

When multiple clients are registered (e.g., `mock`, `litellm`, `openrouter`), use the `run_algorithm` function from `switchyard.libsy` (defined in [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py)) to programmatically select the appropriate backend:

```python
from switchyard.libsy import run_algorithm
from switchyard_litellm.client import LiteLLMSyClient

# Instantiate available clients

clients = {
    "openrouter": LiteLLMSyClient(base_url="http://localhost:8000"),
    "echo": EchoClient()  # Example alternative

}

# Algorithm determines which client to use based on model/routing rules

algorithm = load_routing_rules()

# Execute routing - matches pattern shown in tests/test_libsy_minimal_bindings.py (lines ~57-69)

outcome, metadata = await run_algorithm(algorithm, clients)

```

The `clients` dictionary maps string identifiers to objects implementing the async `call` method. The routing algorithm returns the selected client's response alongside metadata describing the routing decision.

## Key Configuration Parameters

When configuring OpenRouter connectivity, these parameters control the client behavior:

- **`base_url`**: The URL of your LiteLLM gateway or direct OpenRouter-compatible endpoint (e.g., `http://localhost:8000`).
- **`api_key_env`**: The name of the environment variable containing your OpenRouter API key (default: `OPENROUTER_API_KEY`).
- **Model naming**: Use OpenRouter's full model identifiers (e.g., `openrouter/moonshotai/kimi-k3`) in the request's `model` field to ensure proper routing through OpenRouter's catalog.

The implementation in [`examples/experimental/litellm/src/switchyard_litellm/client.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/experimental/litellm/src/switchyard_litellm/client.py) automatically handles LiteLLM's `acompletion` calls and translates `ModelResponse` objects back into Switchyard's standardized `LlmResponse` format.

## Summary

- **Direct instantiation**: Import `LiteLLMSyClient` from `switchyard_litellm.client` and call `await client.call(request)` for standalone Python usage.
- **Server configuration**: Define `[llm_clients.openrouter]` in a TOML file to auto-wire the client at startup, reading credentials from the `OPENROUTER_API_KEY` environment variable.
- **Integration point**: The `call` method in `LiteLLMSyClient` (line ~319) serves as the bridge between Switchyard's Rust routing core and OpenRouter's HTTP API.
- **Request normalization**: Helper functions at lines ~17-95 convert internal `_messages` and `_tools` formats to LiteLLM/OpenRouter schemas.

## Frequently Asked Questions

### What environment variable stores the OpenRouter API key?

Switchyard expects the `OPENROUTER_API_KEY` environment variable by default. When using server-side TOML configuration, the `api_key_env` field under `[llm_clients.openrouter]` specifies this variable name, allowing flexibility if your deployment uses different naming conventions.

### Can I use OpenRouter without the LiteLLM gateway?

While the reference implementation in [`examples/experimental/litellm/src/switchyard_litellm/client.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/experimental/litellm/src/switchyard_litellm/client.py) uses LiteLLM as an intermediary, you could implement a custom client class conforming to the async `call(request)` interface that speaks directly to OpenRouter's REST API. The `LiteLLMSyClient` is simply the provided reference implementation that handles the translation layer.

### How does Switchyard handle context window errors from OpenRouter?

The `LiteLLMSyClient` captures LiteLLM-specific exceptions like `ContextWindowExceededError` and translates them into Switchyard's internal error representations. This allows the routing engine to potentially fall back to alternative clients or models when OpenRouter reports token limit violations, maintaining service availability according to your routing rules.

### Where is the client implementation located?

The primary OpenRouter-compatible client implementation resides in [`examples/experimental/litellm/src/switchyard_litellm/client.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/experimental/litellm/src/switchyard_litellm/client.py) within the NVIDIA-NeMo/Switchyard repository. This file contains the `LiteLLMSyClient` class definition and the conversion helpers for message formatting. Configuration examples appear in [`benchmark/server-configs/tb-lite-single-opus-4-7.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/server-configs/tb-lite-single-opus-4-7.toml).