# How switchyard‑llm‑client Enables Upstream Model Calls in NVIDIA Switchyard

> Discover how switchyard-llm-client acts as an HTTP bridge to make authenticated, protocol-specific calls to upstream LLM models like OpenAI, Anthropic, and NVIDIA NIM.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**switchyard‑llm‑client serves as the core HTTP bridge that transforms provider‑agnostic LLM requests into authenticated, protocol‑specific calls to upstream models like OpenAI, Anthropic, and NVIDIA NIM.**

The `switchyard‑llm‑client` crate is the network execution layer within the NVIDIA‑NeMo/Switchyard ecosystem. It consumes neutral intermediate representations (IR) of LLM requests generated by higher‑level routing logic and translates them into concrete HTTP POST requests against provider‑specific endpoints. All actual network traffic, retry logic, and authentication handling are encapsulated within this component, allowing the rest of the Switchyard framework to remain agnostic to upstream API differences.

## Core Architecture of switchyard‑llm‑client

The public API centers on the **`TranslatingLlmClient`** struct, defined in [[`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs). This struct acts as the single entry point for executing model calls.

When initialized via `TranslatingLlmClient::new()`, the client performs several setup tasks:

- Validates a list of **`ModelConfig`** objects, each describing a model name, its default backend, and optional alternate backends for different wire formats.
- Constructs a shared **`reqwest::Client`** for connection pooling.
- Builds an internal map from model IDs to their respective **`ModelConfig`** instances for fast lookup during request processing.

## How switchyard‑llm‑client Translates Requests to Upstream Calls

During execution, the client transforms abstract LLM requests into concrete HTTP traffic through a six‑stage pipeline:

### Backend Selection

The method `backend_for(&model, format)` determines which **`Backend`** instance to use for a given request. It selects the default backend or an alternate one based on the requested **`WireFormat`**, enabling the same model to be called via different protocols (e.g., OpenAI‑compatible vs. native Anthropic) depending on routing needs.

### Request Encoding

Using the **`switchyard‑translation`** crate, the neutral `LlmRequest` structure is serialized into the backend’s specific wire format. This translation layer handles differences in message schemas, ensuring that a single internal representation can generate valid JSON for OpenAI, Anthropic Messages, or NVIDIA NIM formats.

### Authentication and Header Management

Before transmission, the client sanitizes reserved headers—including `authorization` and `x-api-key`—preventing credential leakage from upstream request metadata. It then injects backend‑specific static headers and dynamic credentials (such as bearer tokens or API keys) required by the target provider.

### HTTP Transport with Retry Logic

The `send_encoded` method (invoked by `call_rewrite_model` and other high‑level interfaces) executes the actual HTTP POST. It implements robust retry logic with exponential backoff, respecting `Retry-After` headers from upstream providers. Retry constants and backoff parameters are defined within [[`client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/client.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs), ensuring configurable resilience against transient failures.

### Response Handling

For buffered responses, `switchyard‑llm‑client` reads the complete response body before returning. For streaming scenarios, it returns a **`Stream`** immediately upon successful header receipt, allowing token‑by‑token processing without buffering the entire response. HTTP error codes are mapped to typed **`LlmClientError`** instances; for example, HTTP 400 responses are translated into context‑window overflow errors, providing actionable diagnostic information to upstream callers.

## Token Counting and Extended Capabilities

Beyond standard completion requests, the client exposes provider‑specific utilities through `supports_count_tokens()` and `count_tokens()`. These methods target Anthropic’s token‑counting endpoints, demonstrating how `switchyard‑llm‑client` extends its translation capabilities beyond simple inference calls to support model‑specific metadata operations.

## Configuring switchyard‑llm‑client for Multiple Providers

The following example demonstrates initializing `TranslatingLlmClient` with configurations for both OpenAI GPT‑4 and Anthropic Claude:

```python
from switchyard.libsy import ModelConfig
from switchyard.libsy_llm_client import TranslatingLlmClient, Backend, WireFormat

# Define backends – each knows its endpoint, auth header, and wire format

openai_backend = Backend(
    url="https://api.openai.com/v1/chat/completions",
    auth_header=("Authorization", f"Bearer {os.getenv('OPENAI_API_KEY')}"),
    wire_format=WireFormat.OpenAiChat,
)

anthropic_backend = Backend(
    url="https://api.anthropic.com/v1/messages",
    auth_header=("x-api-key", os.getenv("ANTHROPIC_API_KEY")),
    wire_format=WireFormat.AnthropicMessages,
)

# Build ModelConfig objects

gpt4_cfg = ModelConfig(
    model_name="gpt-4",
    default_backend=openai_backend,
    other_backends=None,
)

claude_cfg = ModelConfig(
    model_name="claude-2",
    default_backend=anthropic_backend,
    other_backends=None,
)

# Initialise the translating client

llm_client = TranslatingLlmClient.new([gpt4_cfg, claude_cfg])

# Issue a request (neutral LLM request)

request = LlmRequest(
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=50,
)
response = await llm_client.call_rewrite_model(
    model="gpt-4",
    request=request,
)
print(response.choices[0].message.content)

```

To utilize Anthropic‑specific token counting:

```python
if llm_client.supports_count_tokens("claude-2"):
    token_info = await llm_client.count_tokens(
        model="claude-2",
        request=some_request,
    )
    print("Input tokens:", token_info["total_tokens"])

```

## Key Source Files in switchyard‑llm‑client

The functionality of `switchyard‑llm‑client` is distributed across these critical files:

| File | Purpose |
|------|---------|
| [[`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs) | Implements `TranslatingLlmClient`, backend selection, request encoding, HTTP transport, retries, and token‑counting helpers. |
| [[`crates/libsy-llm-client/src/backend.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/backend.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/backend.rs) | Defines the `Backend` struct (URL, auth, wire format) and header validation logic. |
| [[`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) | Provides the high‑level `run` function used by libsy algorithms to drive model calls. |
| [[`crates/libsy-llm-client/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md) | Overview of the crate and its role in the overall architecture. |
| [[`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) | Workspace configuration listing `switchyard‑llm‑client` as a member, enabling Python binding integration. |

## Summary

- **switchyard‑llm‑client** is the exclusive network gateway that converts neutral LLM requests into provider‑specific HTTP calls.
- The **`TranslatingLlmClient`** struct manages backend selection, request translation, and error mapping through its `backend_for()` and `send_encoded()` methods.
- Authentication headers are sanitized and injected per‑backend, with reserved headers like `authorization` explicitly stripped to prevent leakage.
- Retry logic implements exponential backoff with `Retry-After` support, defined in [`client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/client.rs).
- Streaming and buffered response modes are supported, with errors mapped to typed **`LlmClientError`** variants.
- Extended capabilities like Anthropic token counting are exposed via `count_tokens()` and `supports_count_tokens()`.

## Frequently Asked Questions

### How does switchyard‑llm‑client handle authentication for different upstream providers?

According to the source code in [`backend.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/backend.rs), each `Backend` instance stores provider‑specific authentication headers as tuples (header name, value). Before sending requests, `switchyard‑llm‑client` strips reserved headers like `authorization` and `x-api-key` from incoming metadata, then appends the backend’s configured static and credential headers. This ensures clean, provider‑compliant authentication without credential leakage.

### What retry logic does switchyard‑llm‑client implement for failed upstream calls?

The `send_encoded` method in [`client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/client.rs) implements exponential backoff with configurable constants. It inspects upstream `Retry-After` headers when present and applies the specified delay before retrying transient failures. This logic is encapsulated within the client, meaning higher‑level Switchyard algorithms receive resilient network behavior without implementing their own retry mechanisms.

### Can switchyard‑llm‑client route a single model through different wire formats?

Yes. The `ModelConfig` struct supports both a `default_backend` and `other_backends` map keyed by `WireFormat`. The `backend_for()` method selects the appropriate backend based on the requested format, allowing the same logical model (e.g., a fine‑tuned endpoint) to be called via OpenAI‑compatible JSON or native Anthropic protocols depending on routing requirements.

### How are HTTP errors translated into actionable errors for Switchyard applications?

The client maps HTTP status codes to typed `LlmClientError` variants. For instance, HTTP 400 responses are converted to context‑window overflow errors, while other status codes map to specific retryable or fatal error types. This error taxonomy allows the libsy routing layer to make informed decisions about request rewrites, fallbacks, or circuit breaking.