How to Embed Switchyard's Routing Algorithms in a Custom Application
Switchyard exposes its Rust-based routing engine through switchyard.libsy, which provides Python async iterators that yield Step objects—requiring your application to handle CallModel events by invoking your own LLM client and return responses via call.respond().
Switchyard is an open-source LLM routing framework developed by NVIDIA. The codebase centers on libsy, a high-performance Rust library that implements routing algorithms ranging from simple random selection to sophisticated LLM-based classifiers. By embedding these algorithms into your own async Python services, you can leverage Switchyard's decision-making logic while retaining full control over model execution and client integrations.
Understanding the Switchyard Architecture
Switchyard follows a layered architecture that separates the routing logic from the execution environment. At the core is libsy, a Rust crate located in crates/libsy/src/algorithms.rs that defines the Algorithm trait and routing outcomes. The Python façade in switchyard/libsy/__init__.py exposes these Rust structures through PyO3 bindings defined in crates/switchyard-py/src/libsy_bindings.rs.
When you instantiate an algorithm using the factory functions in switchyard/libsy/algorithms.py, you receive an object that implements run_stream(request) → AsyncIterator[Step]. This method drives the routing decision process without blocking, yielding control back to your Python code whenever the algorithm needs external input or has reached a conclusion.
The two primary step types your application must handle are:
- Step.CallModel: Indicates the algorithm requires one or more model candidates to be evaluated.
- Step.Done: Signals that routing is complete and provides the final
selected_model_idalong with an optional cached response.
Implementing the Custom LLM Client
To embed Switchyard's routing logic, you must provide an async-capable LLM client that can fulfill the CallModel requests. The client must implement an interface compatible with the algorithm's expectations, accepting a request mapping and model identifier, and returning a LlmResponse object.
The LlmResponse enum supports both aggregated responses (LlmResponse.Agg) and streaming responses (LlmResponse.Stream), allowing you to integrate with both completion and streaming endpoints from providers like OpenAI, Anthropic, or NVIDIA NIM.
Wiring the Algorithm Stream
The integration pattern follows an async iterator protocol. After creating an algorithm instance using algorithms.random(), algorithms.llm_classifier(), or algorithms.stage_router(), you iterate over algorithm.run_stream(request) using async for. Each iteration yields a Step that requires specific handling.
Handling CallModel Steps
When the iterator yields Step.CallModel(call), your application must execute the LLM call using your client against the candidate models specified in call.models, package the result as a LlmResponse object, and return control to the algorithm via call.respond(response). This step may occur multiple times during a single routing decision, particularly for algorithms that evaluate multiple candidates or perform staged routing.
Processing Done Steps
When the iterator yields Step.Done(outcome), the routing decision is finalized. The outcome object contains selected_model_id (the final model selection), response (an optional cached response if the algorithm already executed the final call), and request (the original request context). If outcome.response is None, your application should invoke the selected model directly to obtain the final completion.
Complete Integration Example
The following example demonstrates a complete implementation using the random routing algorithm with a mock LLM client. This pattern, adapted from examples/libsy.py, shows how to bridge Switchyard's decision engine with your own infrastructure.
import asyncio
from collections.abc import AsyncIterator, Mapping
# Public symbols provided by the Switchyard façade
from switchyard.libsy import LlmResponse, Step, algorithms
# ---------------------------------------------------------
# A tiny LLM client – replace with your own implementation
# ---------------------------------------------------------
class EchoClient:
"""Always returns a fixed completion; useful for demos."""
async def call(
self,
request: Mapping[str, object],
model: str,
) -> LlmResponse.Agg | LlmResponse.Stream:
if request.get("stream"):
async def events() -> AsyncIterator[Mapping[str, object]]:
yield {"preservation": None,
"normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
yield {"preservation": None,
"normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
yield {"preservation": None,
"normalized": [{"MessageStop": {"reason": "end_turn"}}]}
return LlmResponse.Stream(events())
return LlmResponse.Agg({
"model": model,
"outputs": [{"role": "assistant",
"content": [{"type": "text", "text": "Hello"}]}],
})
# ---------------------------------------------------------
# Wiring the algorithm with the client
# ---------------------------------------------------------
async def main() -> None:
# A typical OpenAI‑style request payload
request = {
"model": "auto",
"stream": True,
"messages": [{"role": "user",
"content": [{"type": "text", "text": "Hello"}]}],
}
client = EchoClient()
# Choose any algorithm you need – here we use the random router
# Arguments: list of candidate model IDs, optional weights, optional seed
algorithm = algorithms.random(
["fast", "quality"], # candidate model IDs
weights=[1, 3], # relative probabilities
seed=42, # deterministic for testing
)
# Run the algorithm – it yields Step objects we must handle
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
# Perform the LLM call and give the result back to the algorithm
response = await client.call(call.request, call.models[0])
call.respond(response)
case Step.Done(outcome):
print("🚦 Routing decision:", outcome.selected_model_id)
# If the algorithm already cached a response, use it;
# otherwise you can ask the client again for the selected model.
response = outcome.response or await client.call(
outcome.request,
outcome.selected_model_id,
)
# Inspect the response (Agg or Stream)
match response:
case LlmResponse.Agg(agg):
print("📄 Full response:", agg)
case LlmResponse.Stream(stream):
async for event in stream:
print("🔁 Stream event:", event)
if __name__ == "__main__":
asyncio.run(main())
Available Routing Strategies
The switchyard.libsy.algorithms module provides factory functions for different routing strategies:
- algorithms.random: Probabilistic selection based on configurable weights, implemented in
crates/libsy/src/algorithms/rand.rs. Accepts a list of model IDs, optionalweights, and an optionalseedfor deterministic behavior. - algorithms.llm_classifier: Uses an LLM to classify and route requests based on content analysis.
- algorithms.stage_router: Multi-stage routing that can chain different selection strategies.
Each factory accepts strategy-specific parameters, allowing fine-tuned control over the selection process without modifying the core Rust implementation in crates/libsy/src/algorithms.rs.
Summary
Embedding Switchyard's routing algorithms requires understanding the async iterator pattern and the separation between decision logic and execution:
- Import routing factories from
switchyard.libsy.algorithmsto instantiate Rust-backed algorithms. - Implement an async LLM client that can handle both streaming and aggregated responses.
- Iterate over
algorithm.run_stream()and handleStep.CallModelby invoking your client and callingcall.respond(). - Process
Step.Doneto obtain the finalselected_model_idand any cached response. - Reference the end-to-end example in
examples/libsy.pyfor production-ready patterns.
Frequently Asked Questions
What Python version is required to use switchyard.libsy?
Switchyard requires Python 3.10 or later, as the integration examples utilize structural pattern matching (match/case syntax) introduced in PEP 634. The underlying PyO3 bindings in crates/switchyard-py/src/libsy_bindings.rs target the CPython stable ABI for compatibility with modern Python versions.
Can I use synchronous LLM clients with Switchyard's routing algorithms?
No, the Algorithm.run_stream() method is strictly async and returns an AsyncIterator[Step]. Your LLM client must provide async methods (using async def) to avoid blocking the event loop when handling Step.CallModel events. If your existing client is synchronous, wrap it with asyncio.to_thread() or similar executor patterns before passing responses to call.respond().
How does Step.CallModel differ from Step.Done in the routing lifecycle?
Step.CallModel represents an intermediate state where the algorithm has identified candidate models but needs you to execute the actual inference and return the results for evaluation. Step.Done represents the terminal state where the algorithm has finalized its decision and provides the RoutingOutcome containing the selected_model_id. A single routing pass may yield multiple CallModel steps before reaching Done.
Where is the actual routing logic implemented—Python or Rust?
The core routing logic is implemented in Rust within the libsy crate, specifically in files like crates/libsy/src/algorithms/rand.rs for random selection and crates/libsy/src/algorithms.rs for the trait definitions. The Python package switchyard.libsy provides thin wrappers via PyO3 that expose these Rust algorithms as Python objects, ensuring high-performance decision-making while allowing Python applications to control the I/O boundary.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →