Switchyard Library APIs: Server Management and LLM Routing Reference
The Switchyard library exposes its core functionality through two primary Python packages: switchyard_rust.server for managing the native Rust HTTP server lifecycle, and switchyard_rust.libsy (conveniently re-exported as switchyard.libsy) for constructing routing algorithms including LLM classifiers and stage routers.
Switchyard is a dual-language library developed by NVIDIA that lets Python applications orchestrate Large Language Model traffic while delegating high-performance routing and serving logic to a native Rust implementation. Mastering the Switchyard library APIs enables developers to programmatically control server initialization, manage model endpoints, and implement sophisticated routing strategies for intelligent traffic distribution.
Server API: Managing the Native Rust Runtime
The switchyard_rust.server module provides the primary interface for starting and managing the native Switchyard server. The Server class acts as a thin Python wrapper around the compiled Rust binary, loading the native library on demand and providing a Pythonic context manager for lifecycle management.
Starting the Server with Configuration
Instantiate the Server class by passing a TOML deployment configuration (or file path) and optional port parameter. Setting port=0 instructs the operating system to select an available ephemeral port automatically.
from switchyard_rust.server import Server
# Load TOML configuration and start the server
with Server("routes.toml", port=0) as server:
print(f"Running on http://{server.base_url}")
# Client code can now call OpenAI-compatible endpoints
According to the source code in switchyard_rust/server.py, the Server class handles native library loading and exposes the OpenAI/Anthropic-compatible HTTP endpoints through the underlying Rust binary.
Server Instance Methods and Properties
The Server instance provides several key attributes and methods for runtime inspection and control:
port– Returns the actual port number the server is listening on (particularly useful whenport=0was specified).base_url– Provides a convenient base URL string for constructing client requests.caller_auth_kind(model)– Returns the required authentication kind for a specific model name.close(timeout_secs)– Initiates a graceful shutdown of the server with the specified timeout.
Algorithm API: Implementing Routing Logic
The switchyard_rust.libsy module contains the core routing functionality executed by the Rust libsy crate. All routing algorithms implement the Algorithm interface, which defines the contract for processing LLM requests and determining target models.
The Algorithm Class and run_stream Method
Every algorithm instance implements the run_stream method, which returns an async iterator yielding routing decisions. The method signature accepts request metadata and optional headers:
algorithm.run_stream(
request: Mapping[str, object],
headers: Mapping[str, str] | None = None
) -> AsyncIterator[Step.CallModel | Step.Done]
As implemented in switchyard_rust/libsy.py, the iterator yields:
Step.CallModel– Indicates which model(s) to invoke for the current request.Step.Done– Signifies the final routing outcome with the selected model ID.
Routing Algorithms and Classifiers
The Switchyard library APIs provide several factory functions for constructing routing algorithms tailored to different operational requirements.
Capability Classifier
The capability classifier selects a target based on task capability predictions, routing to an efficient model for simple tasks and a capable model for complex ones. Configure it using LlmClassifierConfig.capability():
from switchyard.libsy import llm_classifier, LlmClassifierConfig, TaskClassifierConfig
cfg = LlmClassifierConfig.capability(
judge_target="classifier",
efficient_target="gpt-3.5",
capable_target="gpt-4",
config=TaskClassifierConfig(
base_threshold=0.5,
session_affinity=True,
),
)
alg = llm_classifier(cfg)
Escalation Classifier
The escalation classifier first attempts to use an efficient model, then escalates to a capable model if the response suggests insufficient performance. Configuration uses LlmClassifierConfig.escalation():
from switchyard.libsy import llm_classifier, LlmClassifierConfig, EscalationClassifierConfig
cfg = LlmClassifierConfig.escalation(
judge_target="classifier",
efficient_target="gpt-3.5",
capable_target="gpt-4",
config=EscalationClassifierConfig(confirmations=2, recent_turn_window=28),
)
alg = llm_classifier(cfg)
Custom Classifier
For domain-specific routing, the custom classifier routes among user-defined targets using JSON-schema-validated labels:
from switchyard.libsy import llm_classifier, LlmClassifierConfig, CustomClassifierConfig
cfg = LlmClassifierConfig.custom(
judge_target="classifier",
targets=[("fast", "gpt-3.5"), ("smart", "gpt-4")],
default_target="fast",
config=CustomClassifierConfig(
prompt="Classify the task",
response_schema={"type": "string", "enum": ["fast", "smart"]},
selector="label",
),
)
alg = llm_classifier(cfg)
All classifier configurations are defined in switchyard_rust/libsy.py, which provides the binding layer to the Rust implementation.
Stage Router API
The stage router algorithm determines when to switch from an efficient target to a capable one based on classifier confidence scores. The stage_router factory function creates this algorithm:
from switchyard.libsy import stage_router
alg = stage_router(
capable_target="gpt-4",
efficient_target="gpt-3.5",
picker="max_confidence",
confidence_threshold=0.8,
recent_window=10,
escalation_note="Escalated due to low confidence",
)
Utility Algorithms
For testing and experimental workloads, switchyard_rust/libsy.py provides additional helpers:
noop()– Returns a no-op algorithm that immediately returns a deterministic outcome without invoking any models.random(targets, weights=None, seed=None)– Selects targets uniformly or with specified weights for A/B testing scenarios.
Import Convenience Layer
To simplify client code, the switchyard.libsy.algorithms module re-exports all factory functions from the underlying Rust bindings. This allows users to import directly from the top-level switchyard package rather than referencing the internal switchyard_rust structure:
from switchyard.libsy import llm_classifier, stage_router, LlmClassifierConfig
The re-export definitions reside in switchyard/libsy/algorithms.py, which serves as the canonical import location for Python applications.
End-to-End Integration Example
The following example demonstrates the complete workflow: starting a server, configuring a capability classifier, and processing a request through the algorithm:
import asyncio
from switchyard_rust.server import Server
from switchyard.libsy import llm_classifier, LlmClassifierConfig, TaskClassifierConfig
async def main():
# Initialize server with TOML configuration
server = Server("examples/routes.toml", port=0)
await server.__aenter__()
try:
# Configure capability-based routing
clf_cfg = LlmClassifierConfig.capability(
judge_target="classifier",
efficient_target="gpt-3.5",
capable_target="gpt-4",
config=TaskClassifierConfig(base_threshold=0.6, session_affinity=True),
)
algorithm = llm_classifier(clf_cfg)
# Execute routing decision
request = {
"model": "gpt-3.5",
"messages": [{"role": "user", "content": "Summarize Shakespeare"}]
}
async for step in algorithm.run_stream(request):
if isinstance(step, Step.CallModel):
print("Calling model:", step.call.models)
else: # Step.Done
print("Final selection:", step.outcome.selected_model_id)
finally:
await server.__aexit__(None, None, None)
asyncio.run(main())
In this implementation, the Server object manages the model lifecycle and HTTP endpoints, while the Algorithm operates independently to determine routing decisions through the run_stream async iterator.
Summary
- Server API (
switchyard_rust/server.py): Use theServerclass to start the native Rust HTTP server, manage ports, and handle authentication configurations through a Pythonic context manager interface. - Algorithm API (
switchyard_rust/libsy.py): Implement routing logic using factory functions likellm_classifier()andstage_router(), which returnAlgorithminstances that yield routing decisions viarun_stream(). - Classifier Types: Choose between capability classifiers (task-based routing), escalation classifiers (retry-based routing), and custom classifiers (schema-validated routing) based on your traffic distribution requirements.
- Convenience Imports: Import from
switchyard.libsyrather thanswitchyard_rust.libsyfor cleaner dependency management in production applications.
Frequently Asked Questions
What is the difference between switchyard_rust.server and switchyard.libsy?
switchyard_rust.server manages the HTTP server lifecycle and OpenAPI-compatible endpoints, wrapping the native Rust binary in a Python class. switchyard.libsy (re-exported from switchyard_rust.libsy) provides the routing algorithm interface for determining which LLM should handle specific requests. The server manages how requests are received, while the algorithms determine where requests are sent.
How do I configure a custom routing algorithm in Switchyard?
Import LlmClassifierConfig.custom() from switchyard.libsy to define targets with custom labels and JSON schemas. Specify your targets as tuples of (label, model_name), provide a validation schema, and set a default target for fallbacks. The classifier will validate responses against your schema and route accordingly.
Can I run the Switchyard server without using the context manager?
Yes, though the context manager (with Server(...) as server:) is recommended for automatic cleanup. You can manually instantiate Server("config.toml", port=0), call await server.__aenter__() to start, and explicitly call await server.__aexit__(None, None, None) or server.close(timeout_secs) when shutting down.
What data does the Algorithm.run_stream method return?
The run_stream method returns an async iterator that yields Step objects from the Rust runtime. Each iteration yields either a Step.CallModel containing the target model identifier and call parameters, or a Step.Done object containing the final routing outcome including the selected model ID. Your application logic must handle these steps to execute the model calls and return responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →