How Python Applications Interact with Switchyard's Routing Algorithms Using switchyard-py
Python applications interact with Switchyard's routing algorithms by importing factory functions from the switchyard-py package (distributed as nemo-switchyard), which exposes Rust-implemented algorithms through PyO3 bindings in switchyard_rust/libsy.py.
Switchyard is an intelligent routing layer for LLM workloads developed by NVIDIA. The switchyard-py Python bindings allow developers to leverage high-performance routing algorithms—implemented in the Rust libsy crate—without writing any Rust code, enabling seamless integration into existing async Python applications.
Architecture of the Python-Rust Bridge
Switchyard's Python integration relies on a three-tier architecture that isolates the compiled Rust logic from user-facing Python APIs.
crates/libsy– The core Rust crate implementing routing algorithms (random selection, LLM-based classification, stage routers) and the streaming protocol.switchyard_rust/libsy.py– A Python façade that lazily loads the compiled native library viaload_native()and re-exports Rust symbols (Algorithm,Step,LlmResponse). It uses__getattr__to forward calls to the native library while providing type stubs for static analysis.switchyard/libsy/algorithms.py– A convenience wrapper that imports factory functions (random,llm_classifier,stage_router) fromswitchyard_rust.libsyand exposes them in the publicswitchyard.libsy.algorithmsnamespace.
This design ensures that all heavy computational work occurs in optimized Rust code, while Python handles orchestration and I/O.
Installing switchyard-py
Install the package via pip using its distribution name:
pip install nemo-switchyard
The package installs the compiled Rust extensions alongside the Python façade modules.
Using Routing Algorithms in Python
Importing the Algorithm Factories
Begin by importing the algorithms module, which provides factory functions for creating router instances:
from switchyard.libsy import algorithms
This import path resolves through switchyard/libsy/algorithms.py, which ultimately loads the Rust-backed implementations via switchyard_rust/libsy.py.
Creating a Routing Algorithm Instance
Instantiate a specific router by calling one of the factory functions. For example, create a random router that selects between targets "fast" and "quality" with weighted probabilities:
algorithm = algorithms.random(
["fast", "quality"],
weights=[1, 3],
seed=42
)
The returned object is an Algorithm instance backed by the Rust implementation.
Running the Algorithm Stream
Execute the routing logic by calling run_stream(), which returns an async iterator of Step objects:
async for step in algorithm.run_stream(request):
...
The request parameter is a dictionary containing the LLM request payload (e.g., {"model": "auto", "messages": [...]}).
Handling Step Types
The async iterator yields Step objects representing distinct phases of the routing lifecycle:
Step.CallModel– Contains aModelCallwith the original request, a list of candidate models, and arespondcallback. Your application must call the specified model and forward the response viacall.respond().Step.Done– Contains aRoutingOutcomeexposingselected_model_id,fallback_models, and the finalresponse(either an aggregate or stream).
Complete Implementation Examples
Random Router with Echo Client
The following example from examples/libsy.py demonstrates a complete integration using a mock Echo client:
#!/usr/bin/env python3
import asyncio
from collections.abc import AsyncIterator, Mapping
from switchyard.libsy import LlmResponse, Step, algorithms
class EchoClient:
"""Returns a fixed completion for any selected target."""
async def call(self, request: Mapping[str, object], model: str) -> LlmResponse.Agg | LlmResponse.Stream:
if request.get("stream"):
async def events() -> AsyncIterator[Mapping[str, object]]:
yield {"preservation": None, "normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
yield {"preservation": None, "normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
yield {"preservation": None, "normalized": [{"MessageStop": {"reason": "end_turn"}}]}
return LlmResponse.Stream(events())
return LlmResponse.Agg(
{
"model": model,
"outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}],
}
)
async def main() -> None:
request = {
"model": "auto",
"stream": True,
"messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}],
}
client = EchoClient()
# Build a random router that chooses between two targets.
algorithm = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
# Forward the request to the selected model via the client.
call.respond(await client.call(call.request, call.models[0]))
case Step.Done(outcome):
print("Decision:", outcome.selected_model_id)
# The outcome may already contain a response; otherwise we ask the client again.
response = outcome.response or await client.call(outcome.request, outcome.selected_model_id)
match response:
case LlmResponse.Agg(agg):
print("Response:", agg)
case LlmResponse.Stream(stream):
async for ev in stream:
print("Response event:", ev)
if __name__ == "__main__":
asyncio.run(main())
LLM-Based Classifier
Configure a classifier router that uses an LLM to determine task capability:
from switchyard.libsy import algorithms, LlmResponse, Step
config = algorithms.llm_classifier.capability(
judge_target="judge",
efficient_target="fast",
capable_target="quality",
config=algorithms.TaskClassifierConfig(
base_threshold=0.7,
session_affinity=False,
recent_turn_window=10,
max_output_tokens=1024,
prompt="Classify the user request",
),
)
algorithm = algorithms.llm_classifier(config)
async for step in algorithm.run_stream(your_request):
# Handle Step.CallModel and Step.Done as shown in the random router example
pass
The TaskClassifierConfig class is defined in the type stubs within switchyard_rust/libsy.py.
Stage Router with Escalation
Implement a cascading router that escalates from an efficient model to a capable model based on confidence thresholds:
algorithm = algorithms.stage_router(
capable_target="quality",
efficient_target="fast",
picker="confidence",
confidence_threshold=0.8,
recent_window=20,
escalation_note="Escalating due to low confidence",
deescalation_note="De‑escalating after successful response",
)
This algorithm first attempts routing to the efficient_target, then escalates to capable_target if confidence falls below the specified threshold.
Key Source Files and API Surface
| File | Purpose |
|---|---|
switchyard_rust/libsy.py |
Python façade that lazily loads the compiled Rust libsy library and re-exports symbols (Algorithm, Step, LlmResponse). |
switchyard/libsy/algorithms.py |
Public factory module providing random(), llm_classifier(), and stage_router() functions. |
examples/libsy.py |
End-to-end reference implementation showing algorithm initialization, mock client integration, and step handling. |
docs/core_concepts.md |
Architectural documentation explaining the client-router-backend data flow. |
Summary
- switchyard-py (pip package
nemo-switchyard) exposes Switchyard's Rust routing algorithms to Python via PyO3 bindings inswitchyard_rust/libsy.py. - Applications interact with algorithms asynchronously using
run_stream(), which yieldsStep.CallModelandStep.Doneobjects. - The
switchyard.libsy.algorithmsmodule provides factory functions (random,llm_classifier,stage_router) that instantiate Rust-backedAlgorithmobjects. - Users implement client logic to fulfill
CallModelrequests and feed responses back via therespond()callback. - All heavy routing logic executes in the compiled Rust
libsycrate, while Python manages async I/O and business logic.
Frequently Asked Questions
What is switchyard-py and how does it relate to the Switchyard project?
switchyard-py is the official Python client library for NVIDIA's Switchyard routing framework. It wraps the Rust libsy crate using PyO3 bindings, allowing Python applications to invoke Switchyard's routing algorithms without requiring Rust development expertise or toolchains.
Do I need to write Rust code to use Switchyard's routing algorithms?
No. The switchyard-py package distributes pre-compiled Rust extensions. You interact with the algorithms entirely through Python imports from switchyard.libsy and switchyard.libsy.algorithms, as demonstrated in examples/libsy.py.
How do I handle streaming responses when using switchyard-py?
The run_stream() method returns an async iterator of Step objects. When you encounter LlmResponse.Stream (either in a CallModel step or the final Done outcome), iterate over its async generator to receive streaming chunks. For non-streaming use cases, handle LlmResponse.Agg to receive complete responses.
Where are the routing algorithm implementations actually located?
The core algorithms reside in the Rust crate crates/libsy. The Python files (switchyard_rust/libsy.py and switchyard/libsy/algorithms.py) provide only thin façades and type stubs that forward calls to the compiled native library loaded at runtime.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →