How to Use Switchyard for A/B Traffic Splitting with LLMs: A Complete Guide

Switchyard enables zero-code A/B traffic splitting by routing requests between Large Language Model (LLM) backends based on configurable algorithms that require no changes to client application code.

Switchyard is a Python-first orchestration layer developed by NVIDIA that sits between client applications and LLM backends. According to the NVIDIA-NeMo/Switchyard source code, the framework's core responsibility is routing: dynamically determining which backend services a given request using pluggable algorithms. This architecture makes Switchyard ideal for running controlled A/B experiments, canary deployments, and traffic splitting across different model versions or providers.

Understanding Switchyard's Routing Architecture

Switchyard implements a hybrid Python-Rust architecture designed for high-performance request distribution.

Core Components

The framework consists of four primary architectural layers:

  • Python façade (switchyard.libsy): Exposes routing algorithms as Python callables in switchyard/libsy/algorithms.py. This module provides the high-level interface for configuring traffic splits.

  • Rust implementation (switchyard_rust.libsy): Provides performance-critical routing primitives written in Rust, including random selection, stage-based routing, and LLM-based classification. The bridge logic resides in switchyard_rust/libsy.py.

  • Server component (switchyard_rust.server): Hosts an HTTP endpoint in switchyard_rust/server.py that mimics the OpenAI and Anthropic API specifications. All incoming requests pass through the routing layer before reaching target LLMs.

  • Benchmarking utilities: The scripts/benchmark_routing_algorithms.py script demonstrates how to validate traffic distribution accuracy and measure routing overhead.

Routing Algorithms for A/B Testing

Switchyard supports multiple algorithms for traffic splitting, configured via routing profiles (Python dictionaries or TOML files):

Algorithm Mechanism Best Use Case
random Uniform or weighted selection across targets Simple percentage-based splits (e.g., 50/50, 90/10)
stage_router Routes based on a stage field in request metadata Deterministic splits when clients can label requests as "control" or "treatment"
llm_classifier Uses an LLM to classify prompt content and route accordingly Semantic A/B testing where routing depends on prompt characteristics

When Switchyard receives a request, the selected algorithm inspects metadata and selects a named target (e.g., model_a or model_b), then forwards the request to the appropriate downstream endpoint.

Configuring A/B Traffic Splitting in Switchyard

A/B experiments in Switchyard require no client code modifications—only routing profile configuration changes.

Random Weighted Routing

For simple percentage-based splits, use the random algorithm in your routing profile. In switchyard/libsy/algorithms.py, this algorithm accepts weighted targets:

Standard 50/50 split:

routing_profile = {
    "algorithm": "random",
    "targets": [
        {"model": "openai/gpt-4o-mini", "weight": 0.5},
        {"model": "openai/gpt-4-turbo", "weight": 0.5},
    ],
}

Weighted canary deployment (80% control, 20% experiment):

routing_profile = {
    "algorithm": "random",
    "targets": [
        {"model": "openai/gpt-4o-mini", "weight": 0.8},
        {"model": "openai/gpt-4-turbo", "weight": 0.2},
    ],
}

Deterministic Stage-Based Routing

For experiments requiring consistent routing based on cohort assignment, use the stage_router algorithm. As implemented in switchyard_rust/libsy.py, this routes based on an explicit stage identifier:

routing_profile = {
    "algorithm": "stage_router",
    "targets": [
        {"model": "openai/gpt-4o-mini", "stage": "control"},
        {"model": "openai/gpt-4-turbo", "stage": "treatment"},
    ],
}

Clients must include the stage field in request metadata for this approach.

Advanced LLM Classifier Routing

For semantic traffic splitting—where routing decisions depend on prompt content—configure the llm_classifier or llm_task_classifier algorithms. These use an auxiliary LLM to categorize incoming prompts and route to appropriate targets based on classification results.

Implementing the Request Flow

The complete request lifecycle through Switchyard for A/B traffic splitting follows this sequence:

  1. Client submission: The application sends a standard OpenAI-compatible request (e.g., chat.completions.create) to the Switchyard endpoint.

  2. Python façade processing: switchyard.libsy receives the request and forwards it to the Rust routing layer (switchyard_rust.libsy).

  3. Algorithm execution: The configured routing algorithm selects a target based on the active profile and request metadata.

  4. Backend forwarding: switchyard_rust/server.py rewrites the request to point to the chosen backend's endpoint (e.g., OpenAI, Anthropic, or local vLLM instance).

  5. Response return: The downstream LLM processes the request and returns the response, which Switchyard passes back to the client unchanged.

The client remains agnostic to which specific model handled the request. Only the routing profile determines the traffic split.

Validating Your A/B Test Configuration

Switchyard provides built-in observability to verify traffic distribution accuracy.

When starting the server, enable routing statistics collection:

switchyard-server --config routes.toml --routing-profile-json a_b.json --routing-stats-json stats.json

After running your workload, inspect the generated stats.json file to confirm:

  • Per-target call counts match your intended percentages
  • Latency distributions across variants
  • Error rates by backend

The scripts/benchmark_routing_algorithms.py utility automates this validation, generating detailed reports on split fidelity and routing performance.

For integration testing, reference tests/test_libsy_minimal_bindings.py, which verifies that routing works correctly with streamed responses—a critical check when validating real-time LLM interactions.

Summary

  • Switchyard enables A/B traffic splitting with LLMs through configurable routing profiles that require no client code changes.
  • The architecture combines a Python façade (switchyard/libsy/algorithms.py) with a high-performance Rust backend (switchyard_rust/libsy.py) for routing decisions.
  • Random weighted routing handles percentage-based splits, while stage_router provides deterministic cohort assignment.
  • Routing profiles are plain Python dictionaries specifying algorithms and target weights, passed to the client or server via CLI flags.
  • Built-in observability tools (--routing-stats-json) validate that actual traffic splits match experimental design parameters.

Frequently Asked Questions

How does Switchyard handle traffic splitting without modifying client code?

Switchyard acts as a transparent proxy. Clients send standard OpenAI API requests to the Switchyard server endpoint instead of directly to the LLM provider. The server applies the configured routing algorithm from switchyard/libsy/algorithms.py to select the target backend, then forwards the request. The client receives the response as if it came from a single model, maintaining full API compatibility while enabling backend experimentation.

What is the performance overhead of Switchyard's routing layer?

The routing logic is implemented in Rust (switchyard_rust/libsy.py) with Python bindings, providing microsecond-level latency for routing decisions. According to the scripts/benchmark_routing_algorithms.py implementation, the overhead is negligible compared to LLM inference latency, though you should benchmark your specific routing profile using the provided utilities to validate performance characteristics for your traffic volume.

Can I use Switchyard for multi-armed bandit experiments or dynamic traffic allocation?

While Switchyard's core algorithms (random, stage_router, llm_classifier) support static A/B configurations, the routing profile system is extensible. You can implement custom routing logic in Python by extending the algorithms in switchyard/libsy/algorithms.py or by dynamically updating the routing profile JSON based on external metrics. However, the current stable implementation focuses on static or classifier-based routing rather than real-time adaptive allocation.

How do I ensure my A/B test maintains user session consistency?

For session-aware routing—where the same user must consistently hit the same model variant—use the stage_router algorithm with deterministic stage assignment. Assign users to "control" or "treatment" stages based on user ID hashing or session tokens in your application layer, then pass the appropriate stage value in request metadata. This ensures consistent routing without requiring sticky sessions in the Switchyard server itself.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →