# How Switchyard Routes LLM Requests Across Different Providers

> Discover how Switchyard routes LLM requests across providers using its Python-Rust architecture. Learn about configurable algorithms for intelligent target selection and delegation.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Switchyard routes LLM requests across different providers using a hybrid Python-Rust architecture where configurable algorithms in the `libsy` crate select targets based on model names, tasks, or custom strategies, then delegate to provider-specific HTTP clients.**

Switchyard, NVIDIA's open-source LLM serving platform, abstracts away provider differences through a sophisticated routing layer. This system inspects incoming requests and directs them to the appropriate backend—whether OpenAI, Anthropic, NVIDIA, or OpenRouter—using algorithms defined in the `libsy` crate and exposed via `switchyard_rust` bindings.

## The Hybrid Architecture: Python Frontend to Rust Core

The routing system spans two languages for performance and ergonomics. **Python wrappers** in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) provide the user-facing API, while the **Rust `libsy` crate** handles the actual routing logic. The `switchyard_rust` bindings bridge these layers, allowing Python code to invoke high-performance Rust algorithms without overhead.

This architecture separates configuration (handled in Python) from execution (optimized in Rust). When you instantiate a `SwitchyardClient` with a specific router, you're selecting a Rust-implemented strategy that operates on a `RoutingTable` built from your TOML deployment configuration.

## Routing Algorithms in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)

The decision logic resides in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), which exposes four primary strategies as thin Python wrappers around Rust implementations. Each algorithm implements a different selection criteria for mapping requests to provider targets.

### Model-Based Classification with `llm_classifier`

The **`llm_classifier`** algorithm routes requests based on the model identifier string (e.g., `openai/gpt-4`, `anthropic/claude-3`). This is the default strategy, parsing the model name to determine the appropriate provider backend. It performs a lookup in the `RoutingTable` to match the namespace prefix against configured targets.

### Task-Based Routing with `llm_task_classifier`

For deployments where the same model name might handle different operations, **`llm_task_classifier`** inspects the request type—chat, completion, embedding, or moderation—to select the optimal target. This allows fine-grained control over which provider handles specific inference tasks.

### Random and Staged failover Strategies

The **`random`** algorithm selects targets probabilistically, useful for A/B testing or load distribution across equivalent providers. For complex deployments, **`stage_router`** composes multiple strategies into a pipeline, attempting a primary router first and falling back to secondary options if the primary returns no valid target.

## Provider Configuration and the Routing Table

Targets are defined in a **TOML deployment file** that the native server reads at startup. Each entry specifies the provider type, model identifier, credentials, and endpoint URLs. The server implementation in [`crates/switchyard-server/src/deployment.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/deployment.rs) (implicitly handling the TOML parsing) constructs a `RoutingTable` from these definitions.

The `RoutingTable` serves as the lookup source for all routing algorithms. When Rust code in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) invokes a routing decision, it passes this immutable table to the selected algorithm, which returns a **target ID** corresponding to the chosen provider configuration.

## The Request Journey: From Client to Provider

Understanding the exact flow reveals how Switchyard achieves seamless provider abstraction:

1. The Python frontend receives a request via `client.chat()` or similar methods
2. The request passes to [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), which calls into the Rust `libsy` router
3. The router invokes the selected algorithm (e.g., `llm_classifier`) against the `RoutingTable`
4. The algorithm returns the target ID for the appropriate provider
5. [`crates/libsy-llm-client/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/lib.rs) translates the generic Switchyard request into the provider-specific HTTP payload format
6. The HTTP client dispatches the request to the actual backend (OpenAI, Anthropic, etc.)

This pipeline ensures that provider-specific quirks—such as differing JSON schemas or authentication headers—are handled in the translation layer rather than forcing users to manage multiple client libraries.

## Observability and Routing Metrics

After dispatch, [`crates/switchyard-server/src/routing_log.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/routing_log.rs) records detailed telemetry about each routing decision. The system tracks **routing overhead**, **model call counts**, and error rates. These metrics can be exported as JSON via the `--routing-stats-json` flag and visualized using built-in reporting scripts.

This logging occurs in the Rust layer for minimal performance impact, capturing the exact target selected and the time spent in routing logic versus actual model inference.

## Code Examples for Custom Routing Strategies

### Default Model-Based Routing

The simplest usage relies on the `llm_classifier` to parse model names and route accordingly:

```python
import asyncio
from switchyard import SwitchyardClient

async def main():
    client = SwitchyardClient()
    response = await client.chat(
        model="openai/gpt-4",
        messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
    )
    print(response["choices"][0]["message"]["content"])

asyncio.run(main())

```

### Random Selection for A/B Testing

To distribute load randomly across configured targets:

```python
from switchyard.libsy import algorithms as routing
from switchyard import SwitchyardClient

# Create a random router

router = routing.random()

# Attach the router to the client

client = SwitchyardClient(router=router)

# The request will be sent to a randomly chosen target defined in the TOML config

await client.chat(model="any-model", messages=[...])

```

### Composed Routing with Fallback

Use `stage_router` to attempt model-based routing first, then fall back to random selection if no specific match exists:

```python
from switchyard.libsy import algorithms as routing

stage = routing.stage_router(
    primary=routing.llm_classifier(),
    fallback=routing.random(),
)

client = SwitchyardClient(router=stage)
await client.chat(model="openai/gpt-4", messages=[...])

```

## Summary

- Switchyard implements routing through a **Rust core** (`libsy` crate) exposed via **Python bindings** in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) for high performance with Pythonic ergonomics.
- Four primary algorithms handle selection: **`llm_classifier`** (by model name), **`llm_task_classifier`** (by task type), **`random`** (probabilistic), and **`stage_router`** (composable pipelines).
- Provider targets are defined in **TOML deployment files** and loaded into an immutable `RoutingTable` parsed by the server implementation.
- The HTTP translation layer in [`crates/libsy-llm-client/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/lib.rs) converts generic requests to provider-specific formats after routing decisions are made.
- Routing decisions and performance metrics are logged via [`crates/switchyard-server/src/routing_log.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/routing_log.rs), supporting JSON export and visualization.

## Frequently Asked Questions

### How does Switchyard decide which LLM provider to use for a request?

Switchyard applies a **configurable routing algorithm** defined in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) that inspects the request model name, task type, or custom criteria. The algorithm looks up the matching target in a `RoutingTable` built from the TOML deployment configuration and returns the specific provider ID to handle the request.

### Can I implement custom routing logic beyond the built-in algorithms?

Yes. While the built-in algorithms in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) cover most use cases, you can compose custom logic using the **`stage_router`** to chain multiple strategies, or extend the Python wrapper interface to interact with the underlying Rust `libsy` routing engine for entirely bespoke selection logic.

### Where are routing decisions logged in the Switchyard codebase?

Routing decisions are recorded in **[`crates/switchyard-server/src/routing_log.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/routing_log.rs)**, which captures the selected target, routing overhead latency, and error states. These metrics can be exported as JSON using the `--routing-stats-json` command-line flag for external analysis and monitoring.

### What file format configures provider targets in Switchyard?

Provider targets are configured using **TOML deployment files** that specify each target's provider type (OpenAI, Anthropic, etc.), model identifiers, and authentication credentials. The server parses these files at startup to construct the `RoutingTable` used by all routing algorithms.