How to Use Switchyard for A/B Traffic Splitting with LLMs: A Complete Guide
Switchyard enables zero-code A/B traffic splitting by routing requests between Large Language Model (LLM) backends based on configurable algorithms that require no changes to client application code.
Switchyard is a Python-first orchestration layer developed by NVIDIA that sits between client applications and LLM backends. According to the NVIDIA-NeMo/Switchyard source code, the framework's core responsibility is routing: dynamically determining which backend services a given request using pluggable algorithms. This architecture makes Switchyard ideal for running controlled A/B experiments, canary deployments, and traffic splitting across different model versions or providers.
Understanding Switchyard's Routing Architecture
Switchyard implements a hybrid Python-Rust architecture designed for high-performance request distribution.
Core Components
The framework consists of four primary architectural layers:
-
Python façade (
switchyard.libsy): Exposes routing algorithms as Python callables inswitchyard/libsy/algorithms.py. This module provides the high-level interface for configuring traffic splits. -
Rust implementation (
switchyard_rust.libsy): Provides performance-critical routing primitives written in Rust, including random selection, stage-based routing, and LLM-based classification. The bridge logic resides inswitchyard_rust/libsy.py. -
Server component (
switchyard_rust.server): Hosts an HTTP endpoint inswitchyard_rust/server.pythat mimics the OpenAI and Anthropic API specifications. All incoming requests pass through the routing layer before reaching target LLMs. -
Benchmarking utilities: The
scripts/benchmark_routing_algorithms.pyscript demonstrates how to validate traffic distribution accuracy and measure routing overhead.
Routing Algorithms for A/B Testing
Switchyard supports multiple algorithms for traffic splitting, configured via routing profiles (Python dictionaries or TOML files):
| Algorithm | Mechanism | Best Use Case |
|---|---|---|
| random | Uniform or weighted selection across targets | Simple percentage-based splits (e.g., 50/50, 90/10) |
| stage_router | Routes based on a stage field in request metadata |
Deterministic splits when clients can label requests as "control" or "treatment" |
| llm_classifier | Uses an LLM to classify prompt content and route accordingly | Semantic A/B testing where routing depends on prompt characteristics |
When Switchyard receives a request, the selected algorithm inspects metadata and selects a named target (e.g., model_a or model_b), then forwards the request to the appropriate downstream endpoint.
Configuring A/B Traffic Splitting in Switchyard
A/B experiments in Switchyard require no client code modifications—only routing profile configuration changes.
Random Weighted Routing
For simple percentage-based splits, use the random algorithm in your routing profile. In switchyard/libsy/algorithms.py, this algorithm accepts weighted targets:
Standard 50/50 split:
routing_profile = {
"algorithm": "random",
"targets": [
{"model": "openai/gpt-4o-mini", "weight": 0.5},
{"model": "openai/gpt-4-turbo", "weight": 0.5},
],
}
Weighted canary deployment (80% control, 20% experiment):
routing_profile = {
"algorithm": "random",
"targets": [
{"model": "openai/gpt-4o-mini", "weight": 0.8},
{"model": "openai/gpt-4-turbo", "weight": 0.2},
],
}
Deterministic Stage-Based Routing
For experiments requiring consistent routing based on cohort assignment, use the stage_router algorithm. As implemented in switchyard_rust/libsy.py, this routes based on an explicit stage identifier:
routing_profile = {
"algorithm": "stage_router",
"targets": [
{"model": "openai/gpt-4o-mini", "stage": "control"},
{"model": "openai/gpt-4-turbo", "stage": "treatment"},
],
}
Clients must include the stage field in request metadata for this approach.
Advanced LLM Classifier Routing
For semantic traffic splitting—where routing decisions depend on prompt content—configure the llm_classifier or llm_task_classifier algorithms. These use an auxiliary LLM to categorize incoming prompts and route to appropriate targets based on classification results.
Implementing the Request Flow
The complete request lifecycle through Switchyard for A/B traffic splitting follows this sequence:
-
Client submission: The application sends a standard OpenAI-compatible request (e.g.,
chat.completions.create) to the Switchyard endpoint. -
Python façade processing:
switchyard.libsyreceives the request and forwards it to the Rust routing layer (switchyard_rust.libsy). -
Algorithm execution: The configured routing algorithm selects a target based on the active profile and request metadata.
-
Backend forwarding:
switchyard_rust/server.pyrewrites the request to point to the chosen backend's endpoint (e.g., OpenAI, Anthropic, or local vLLM instance). -
Response return: The downstream LLM processes the request and returns the response, which Switchyard passes back to the client unchanged.
The client remains agnostic to which specific model handled the request. Only the routing profile determines the traffic split.
Validating Your A/B Test Configuration
Switchyard provides built-in observability to verify traffic distribution accuracy.
When starting the server, enable routing statistics collection:
switchyard-server --config routes.toml --routing-profile-json a_b.json --routing-stats-json stats.json
After running your workload, inspect the generated stats.json file to confirm:
- Per-target call counts match your intended percentages
- Latency distributions across variants
- Error rates by backend
The scripts/benchmark_routing_algorithms.py utility automates this validation, generating detailed reports on split fidelity and routing performance.
For integration testing, reference tests/test_libsy_minimal_bindings.py, which verifies that routing works correctly with streamed responses—a critical check when validating real-time LLM interactions.
Summary
- Switchyard enables A/B traffic splitting with LLMs through configurable routing profiles that require no client code changes.
- The architecture combines a Python façade (
switchyard/libsy/algorithms.py) with a high-performance Rust backend (switchyard_rust/libsy.py) for routing decisions. - Random weighted routing handles percentage-based splits, while stage_router provides deterministic cohort assignment.
- Routing profiles are plain Python dictionaries specifying algorithms and target weights, passed to the client or server via CLI flags.
- Built-in observability tools (
--routing-stats-json) validate that actual traffic splits match experimental design parameters.
Frequently Asked Questions
How does Switchyard handle traffic splitting without modifying client code?
Switchyard acts as a transparent proxy. Clients send standard OpenAI API requests to the Switchyard server endpoint instead of directly to the LLM provider. The server applies the configured routing algorithm from switchyard/libsy/algorithms.py to select the target backend, then forwards the request. The client receives the response as if it came from a single model, maintaining full API compatibility while enabling backend experimentation.
What is the performance overhead of Switchyard's routing layer?
The routing logic is implemented in Rust (switchyard_rust/libsy.py) with Python bindings, providing microsecond-level latency for routing decisions. According to the scripts/benchmark_routing_algorithms.py implementation, the overhead is negligible compared to LLM inference latency, though you should benchmark your specific routing profile using the provided utilities to validate performance characteristics for your traffic volume.
Can I use Switchyard for multi-armed bandit experiments or dynamic traffic allocation?
While Switchyard's core algorithms (random, stage_router, llm_classifier) support static A/B configurations, the routing profile system is extensible. You can implement custom routing logic in Python by extending the algorithms in switchyard/libsy/algorithms.py or by dynamically updating the routing profile JSON based on external metrics. However, the current stable implementation focuses on static or classifier-based routing rather than real-time adaptive allocation.
How do I ensure my A/B test maintains user session consistency?
For session-aware routing—where the same user must consistently hit the same model variant—use the stage_router algorithm with deterministic stage assignment. Assign users to "control" or "treatment" stages based on user ID hashing or session tokens in your application layer, then pass the appropriate stage value in request metadata. This ensures consistent routing without requiring sticky sessions in the Switchyard server itself.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →