How to Configure an LLM Client in Switchyard for OpenRouter

Switchyard routes LLM requests through a configurable client abstraction that translates internal libsy formats into OpenRouter's OpenAI-compatible API, either via direct Python instantiation or TOML-based server configuration.

Switchyard is a high-performance LLM routing engine developed by NVIDIA that separates request routing logic from provider-specific implementations. When you configure an LLM client in Switchyard for OpenRouter, you leverage the LiteLLMSyClient class to handle API translation and credential management. This client conforms to Switchyard's async call contract, enabling seamless integration with the Rust-based routing core while maintaining Python-side flexibility for provider-specific logic.

Architecture of the OpenRouter Client Integration

Switchyard's architecture decouples the HTTP server from the LLM provider through a client registry populated at startup. When routing to OpenRouter, the system chains through three layers: the Switchyard Server receives the request, the libsy routing engine selects the appropriate client, and the LiteLLMSyClient forwards the normalized request to a LiteLLM gateway configured for OpenRouter's endpoint.

The LiteLLMSyClient Implementation

The concrete implementation resides in examples/experimental/litellm/src/switchyard_litellm/client.py. The LiteLLMSyClient class (defined around line 319) implements the required interface for Switchyard's routing engine:

class LiteLLMSyClient:
    async def call(self, request: Mapping[str, Any]) -> LlmResponse:
        ...

This method converts Switchyard's normalized request format—containing fields like _messages, _tools, and _tool_choice—into a LiteLLM-compatible payload. The conversion helpers (lines ~17-95 in the same file) handle the schema mapping, while the client manages OpenRouter-specific headers and error handling for conditions like ContextWindowExceededError.

Configuration-Driven Client Registration

For server deployments, Switchyard supports declarative configuration via TOML files located in benchmark/server-configs/. The [llm_clients.openrouter] section (exemplified in tb-lite-single-opus-4-7.toml, lines ~9-13) defines the endpoint and authentication:

[llm_clients.openrouter]
base_url = "http://localhost:8000"
api_key_env = "OPENROUTER_API_KEY"

At server startup, Switchyard reads the api_key_env value and injects the corresponding environment variable into the client instance, eliminating the need to hardcode credentials in configuration files.

Method 1: Direct Client Instantiation in Python

When operating Switchyard as a library or testing integrations locally, instantiate LiteLLMSyClient directly. Ensure the OPENROUTER_API_KEY environment variable is set, as the client reads this automatically during initialization.

import os
from switchyard_litellm.client import LiteLLMSyClient

# Configure authentication

os.environ["OPENROUTER_API_KEY"] = "sk-your-openrouter-key"

# Initialize client pointing to your LiteLLM gateway

client = LiteLLMSyClient(base_url="http://localhost:8000")

# Construct a libsy-compatible request

request = {
    "model": "openrouter/moonshotai/kimi-k3",
    "messages": [
        {"role": "user", "content": [{"type": "text", "text": "Write a haiku"}]}
    ],
    "max_tokens": 64,
}

# Execute the async call conforming to Switchyard's contract

response = await client.call(request)

print("Model:", response.model)
print("Content:", response.choices[0].message.content)

The call method returns an LlmResponse object containing the model output, token usage, and stop reasons, formatted consistently regardless of the underlying provider.

Method 2: Server-Side TOML Configuration

For production deployments using the Switchyard server binary, define the client entirely through configuration. Create a TOML file that specifies the OpenRouter client parameters:


# my_openrouter_config.toml

[llm_clients.openrouter]
base_url = "http://localhost:8000"   # LiteLLM proxy/gateway URL

api_key_env = "OPENROUTER_API_KEY"   # Environment variable name

Launch the server with this configuration and the API key exposed in your environment:

OPENROUTER_API_KEY=sk-your-key \
switchyard-server --config my_openrouter_config.toml --port 4000

The server instantiates the LiteLLMSyClient automatically, injecting the base_url and resolved API key. All HTTP requests hitting the server are then routed through this configured OpenRouter client.

Routing Requests to the OpenRouter Client

When multiple clients are registered (e.g., mock, litellm, openrouter), use the run_algorithm function from switchyard.libsy (defined in switchyard/libsy/__init__.py) to programmatically select the appropriate backend:

from switchyard.libsy import run_algorithm
from switchyard_litellm.client import LiteLLMSyClient

# Instantiate available clients

clients = {
    "openrouter": LiteLLMSyClient(base_url="http://localhost:8000"),
    "echo": EchoClient()  # Example alternative

}

# Algorithm determines which client to use based on model/routing rules

algorithm = load_routing_rules()

# Execute routing - matches pattern shown in tests/test_libsy_minimal_bindings.py (lines ~57-69)

outcome, metadata = await run_algorithm(algorithm, clients)

The clients dictionary maps string identifiers to objects implementing the async call method. The routing algorithm returns the selected client's response alongside metadata describing the routing decision.

Key Configuration Parameters

When configuring OpenRouter connectivity, these parameters control the client behavior:

  • base_url: The URL of your LiteLLM gateway or direct OpenRouter-compatible endpoint (e.g., http://localhost:8000).
  • api_key_env: The name of the environment variable containing your OpenRouter API key (default: OPENROUTER_API_KEY).
  • Model naming: Use OpenRouter's full model identifiers (e.g., openrouter/moonshotai/kimi-k3) in the request's model field to ensure proper routing through OpenRouter's catalog.

The implementation in examples/experimental/litellm/src/switchyard_litellm/client.py automatically handles LiteLLM's acompletion calls and translates ModelResponse objects back into Switchyard's standardized LlmResponse format.

Summary

  • Direct instantiation: Import LiteLLMSyClient from switchyard_litellm.client and call await client.call(request) for standalone Python usage.
  • Server configuration: Define [llm_clients.openrouter] in a TOML file to auto-wire the client at startup, reading credentials from the OPENROUTER_API_KEY environment variable.
  • Integration point: The call method in LiteLLMSyClient (line ~319) serves as the bridge between Switchyard's Rust routing core and OpenRouter's HTTP API.
  • Request normalization: Helper functions at lines ~17-95 convert internal _messages and _tools formats to LiteLLM/OpenRouter schemas.

Frequently Asked Questions

What environment variable stores the OpenRouter API key?

Switchyard expects the OPENROUTER_API_KEY environment variable by default. When using server-side TOML configuration, the api_key_env field under [llm_clients.openrouter] specifies this variable name, allowing flexibility if your deployment uses different naming conventions.

Can I use OpenRouter without the LiteLLM gateway?

While the reference implementation in examples/experimental/litellm/src/switchyard_litellm/client.py uses LiteLLM as an intermediary, you could implement a custom client class conforming to the async call(request) interface that speaks directly to OpenRouter's REST API. The LiteLLMSyClient is simply the provided reference implementation that handles the translation layer.

How does Switchyard handle context window errors from OpenRouter?

The LiteLLMSyClient captures LiteLLM-specific exceptions like ContextWindowExceededError and translates them into Switchyard's internal error representations. This allows the routing engine to potentially fall back to alternative clients or models when OpenRouter reports token limit violations, maintaining service availability according to your routing rules.

Where is the client implementation located?

The primary OpenRouter-compatible client implementation resides in examples/experimental/litellm/src/switchyard_litellm/client.py within the NVIDIA-NeMo/Switchyard repository. This file contains the LiteLLMSyClient class definition and the conversion helpers for message formatting. Configuration examples appear in benchmark/server-configs/tb-lite-single-opus-4-7.toml.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →