# How to Use Switchyard for A/B Traffic Splitting with LLMs: A Complete Guide

> Master A/B traffic splitting for LLMs with Switchyard. This guide shows you how to route requests between LLM backends effortlessly using zero-code configuration. Enhance your applications today.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard enables zero-code A/B traffic splitting by routing requests between Large Language Model (LLM) backends based on configurable algorithms that require no changes to client application code.**

Switchyard is a Python-first orchestration layer developed by NVIDIA that sits between client applications and LLM backends. According to the NVIDIA-NeMo/Switchyard source code, the framework's core responsibility is **routing**: dynamically determining which backend services a given request using pluggable algorithms. This architecture makes Switchyard ideal for running controlled A/B experiments, canary deployments, and traffic splitting across different model versions or providers.

## Understanding Switchyard's Routing Architecture

Switchyard implements a hybrid Python-Rust architecture designed for high-performance request distribution.

### Core Components

The framework consists of four primary architectural layers:

- **Python façade (`switchyard.libsy`)**: Exposes routing algorithms as Python callables in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py). This module provides the high-level interface for configuring traffic splits.

- **Rust implementation (`switchyard_rust.libsy`)**: Provides performance-critical routing primitives written in Rust, including random selection, stage-based routing, and LLM-based classification. The bridge logic resides in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).

- **Server component (`switchyard_rust.server`)**: Hosts an HTTP endpoint in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py) that mimics the OpenAI and Anthropic API specifications. All incoming requests pass through the routing layer before reaching target LLMs.

- **Benchmarking utilities**: The [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) script demonstrates how to validate traffic distribution accuracy and measure routing overhead.

### Routing Algorithms for A/B Testing

Switchyard supports multiple algorithms for traffic splitting, configured via **routing profiles** (Python dictionaries or TOML files):

| Algorithm | Mechanism | Best Use Case |
|-----------|-----------|---------------|
| **random** | Uniform or weighted selection across targets | Simple percentage-based splits (e.g., 50/50, 90/10) |
| **stage_router** | Routes based on a `stage` field in request metadata | Deterministic splits when clients can label requests as "control" or "treatment" |
| **llm_classifier** | Uses an LLM to classify prompt content and route accordingly | Semantic A/B testing where routing depends on prompt characteristics |

When Switchyard receives a request, the selected algorithm inspects metadata and selects a **named target** (e.g., `model_a` or `model_b`), then forwards the request to the appropriate downstream endpoint.

## Configuring A/B Traffic Splitting in Switchyard

A/B experiments in Switchyard require no client code modifications—only routing profile configuration changes.

### Random Weighted Routing

For simple percentage-based splits, use the `random` algorithm in your routing profile. In [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), this algorithm accepts weighted targets:

**Standard 50/50 split:**

```python
routing_profile = {
    "algorithm": "random",
    "targets": [
        {"model": "openai/gpt-4o-mini", "weight": 0.5},
        {"model": "openai/gpt-4-turbo", "weight": 0.5},
    ],
}

```

**Weighted canary deployment (80% control, 20% experiment):**

```python
routing_profile = {
    "algorithm": "random",
    "targets": [
        {"model": "openai/gpt-4o-mini", "weight": 0.8},
        {"model": "openai/gpt-4-turbo", "weight": 0.2},
    ],
}

```

### Deterministic Stage-Based Routing

For experiments requiring consistent routing based on cohort assignment, use the `stage_router` algorithm. As implemented in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), this routes based on an explicit stage identifier:

```python
routing_profile = {
    "algorithm": "stage_router",
    "targets": [
        {"model": "openai/gpt-4o-mini", "stage": "control"},
        {"model": "openai/gpt-4-turbo", "stage": "treatment"},
    ],
}

```

Clients must include the `stage` field in request metadata for this approach.

### Advanced LLM Classifier Routing

For semantic traffic splitting—where routing decisions depend on prompt content—configure the `llm_classifier` or `llm_task_classifier` algorithms. These use an auxiliary LLM to categorize incoming prompts and route to appropriate targets based on classification results.

## Implementing the Request Flow

The complete request lifecycle through Switchyard for A/B traffic splitting follows this sequence:

1. **Client submission**: The application sends a standard OpenAI-compatible request (e.g., `chat.completions.create`) to the Switchyard endpoint.

2. **Python façade processing**: `switchyard.libsy` receives the request and forwards it to the Rust routing layer (`switchyard_rust.libsy`).

3. **Algorithm execution**: The configured routing algorithm selects a target based on the active profile and request metadata.

4. **Backend forwarding**: [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py) rewrites the request to point to the chosen backend's endpoint (e.g., OpenAI, Anthropic, or local vLLM instance).

5. **Response return**: The downstream LLM processes the request and returns the response, which Switchyard passes back to the client unchanged.

The client remains agnostic to which specific model handled the request. Only the routing profile determines the traffic split.

## Validating Your A/B Test Configuration

Switchyard provides built-in observability to verify traffic distribution accuracy.

When starting the server, enable routing statistics collection:

```bash
switchyard-server --config routes.toml --routing-profile-json a_b.json --routing-stats-json stats.json

```

After running your workload, inspect the generated [`stats.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/stats.json) file to confirm:
- Per-target call counts match your intended percentages
- Latency distributions across variants
- Error rates by backend

The [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) utility automates this validation, generating detailed reports on split fidelity and routing performance.

For integration testing, reference [`tests/test_libsy_minimal_bindings.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_libsy_minimal_bindings.py), which verifies that routing works correctly with streamed responses—a critical check when validating real-time LLM interactions.

## Summary

- Switchyard enables **A/B traffic splitting with LLMs** through configurable routing profiles that require no client code changes.
- The architecture combines a **Python façade** ([`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)) with a **high-performance Rust backend** ([`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)) for routing decisions.
- **Random weighted routing** handles percentage-based splits, while **stage_router** provides deterministic cohort assignment.
- **Routing profiles** are plain Python dictionaries specifying algorithms and target weights, passed to the client or server via CLI flags.
- Built-in **observability tools** (`--routing-stats-json`) validate that actual traffic splits match experimental design parameters.

## Frequently Asked Questions

### How does Switchyard handle traffic splitting without modifying client code?

Switchyard acts as a transparent proxy. Clients send standard OpenAI API requests to the Switchyard server endpoint instead of directly to the LLM provider. The server applies the configured routing algorithm from [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) to select the target backend, then forwards the request. The client receives the response as if it came from a single model, maintaining full API compatibility while enabling backend experimentation.

### What is the performance overhead of Switchyard's routing layer?

The routing logic is implemented in Rust ([`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)) with Python bindings, providing microsecond-level latency for routing decisions. According to the [`scripts/benchmark_routing_algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/benchmark_routing_algorithms.py) implementation, the overhead is negligible compared to LLM inference latency, though you should benchmark your specific routing profile using the provided utilities to validate performance characteristics for your traffic volume.

### Can I use Switchyard for multi-armed bandit experiments or dynamic traffic allocation?

While Switchyard's core algorithms (`random`, `stage_router`, `llm_classifier`) support static A/B configurations, the routing profile system is extensible. You can implement custom routing logic in Python by extending the algorithms in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) or by dynamically updating the routing profile JSON based on external metrics. However, the current stable implementation focuses on static or classifier-based routing rather than real-time adaptive allocation.

### How do I ensure my A/B test maintains user session consistency?

For session-aware routing—where the same user must consistently hit the same model variant—use the `stage_router` algorithm with deterministic stage assignment. Assign users to "control" or "treatment" stages based on user ID hashing or session tokens in your application layer, then pass the appropriate `stage` value in request metadata. This ensures consistent routing without requiring sticky sessions in the Switchyard server itself.