How to Embed Switchyard's Routing Algorithms in a Custom Application

Switchyard exposes its Rust-based routing engine through switchyard.libsy, which provides Python async iterators that yield Step objects—requiring your application to handle CallModel events by invoking your own LLM client and return responses via call.respond().

Switchyard is an open-source LLM routing framework developed by NVIDIA. The codebase centers on libsy, a high-performance Rust library that implements routing algorithms ranging from simple random selection to sophisticated LLM-based classifiers. By embedding these algorithms into your own async Python services, you can leverage Switchyard's decision-making logic while retaining full control over model execution and client integrations.

Understanding the Switchyard Architecture

Switchyard follows a layered architecture that separates the routing logic from the execution environment. At the core is libsy, a Rust crate located in crates/libsy/src/algorithms.rs that defines the Algorithm trait and routing outcomes. The Python façade in switchyard/libsy/__init__.py exposes these Rust structures through PyO3 bindings defined in crates/switchyard-py/src/libsy_bindings.rs.

When you instantiate an algorithm using the factory functions in switchyard/libsy/algorithms.py, you receive an object that implements run_stream(request) → AsyncIterator[Step]. This method drives the routing decision process without blocking, yielding control back to your Python code whenever the algorithm needs external input or has reached a conclusion.

The two primary step types your application must handle are:

  • Step.CallModel: Indicates the algorithm requires one or more model candidates to be evaluated.
  • Step.Done: Signals that routing is complete and provides the final selected_model_id along with an optional cached response.

Implementing the Custom LLM Client

To embed Switchyard's routing logic, you must provide an async-capable LLM client that can fulfill the CallModel requests. The client must implement an interface compatible with the algorithm's expectations, accepting a request mapping and model identifier, and returning a LlmResponse object.

The LlmResponse enum supports both aggregated responses (LlmResponse.Agg) and streaming responses (LlmResponse.Stream), allowing you to integrate with both completion and streaming endpoints from providers like OpenAI, Anthropic, or NVIDIA NIM.

Wiring the Algorithm Stream

The integration pattern follows an async iterator protocol. After creating an algorithm instance using algorithms.random(), algorithms.llm_classifier(), or algorithms.stage_router(), you iterate over algorithm.run_stream(request) using async for. Each iteration yields a Step that requires specific handling.

Handling CallModel Steps

When the iterator yields Step.CallModel(call), your application must execute the LLM call using your client against the candidate models specified in call.models, package the result as a LlmResponse object, and return control to the algorithm via call.respond(response). This step may occur multiple times during a single routing decision, particularly for algorithms that evaluate multiple candidates or perform staged routing.

Processing Done Steps

When the iterator yields Step.Done(outcome), the routing decision is finalized. The outcome object contains selected_model_id (the final model selection), response (an optional cached response if the algorithm already executed the final call), and request (the original request context). If outcome.response is None, your application should invoke the selected model directly to obtain the final completion.

Complete Integration Example

The following example demonstrates a complete implementation using the random routing algorithm with a mock LLM client. This pattern, adapted from examples/libsy.py, shows how to bridge Switchyard's decision engine with your own infrastructure.

import asyncio
from collections.abc import AsyncIterator, Mapping

# Public symbols provided by the Switchyard façade

from switchyard.libsy import LlmResponse, Step, algorithms


# ---------------------------------------------------------

# A tiny LLM client – replace with your own implementation

# ---------------------------------------------------------

class EchoClient:
    """Always returns a fixed completion; useful for demos."""
    async def call(
        self,
        request: Mapping[str, object],
        model: str,
    ) -> LlmResponse.Agg | LlmResponse.Stream:
        if request.get("stream"):
            async def events() -> AsyncIterator[Mapping[str, object]]:
                yield {"preservation": None,
                       "normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
                yield {"preservation": None,
                       "normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
                yield {"preservation": None,
                       "normalized": [{"MessageStop": {"reason": "end_turn"}}]}
            return LlmResponse.Stream(events())

        return LlmResponse.Agg({
            "model": model,
            "outputs": [{"role": "assistant",
                         "content": [{"type": "text", "text": "Hello"}]}],
        })


# ---------------------------------------------------------

# Wiring the algorithm with the client

# ---------------------------------------------------------

async def main() -> None:
    # A typical OpenAI‑style request payload

    request = {
        "model": "auto",
        "stream": True,
        "messages": [{"role": "user",
                      "content": [{"type": "text", "text": "Hello"}]}],
    }

    client = EchoClient()

    # Choose any algorithm you need – here we use the random router

    # Arguments: list of candidate model IDs, optional weights, optional seed

    algorithm = algorithms.random(
        ["fast", "quality"],          # candidate model IDs

        weights=[1, 3],               # relative probabilities

        seed=42,                      # deterministic for testing

    )

    # Run the algorithm – it yields Step objects we must handle

    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Perform the LLM call and give the result back to the algorithm

                response = await client.call(call.request, call.models[0])
                call.respond(response)

            case Step.Done(outcome):
                print("🚦 Routing decision:", outcome.selected_model_id)
                # If the algorithm already cached a response, use it;

                # otherwise you can ask the client again for the selected model.

                response = outcome.response or await client.call(
                    outcome.request,
                    outcome.selected_model_id,
                )
                # Inspect the response (Agg or Stream)

                match response:
                    case LlmResponse.Agg(agg):
                        print("📄 Full response:", agg)
                    case LlmResponse.Stream(stream):
                        async for event in stream:
                            print("🔁 Stream event:", event)


if __name__ == "__main__":
    asyncio.run(main())

Available Routing Strategies

The switchyard.libsy.algorithms module provides factory functions for different routing strategies:

  • algorithms.random: Probabilistic selection based on configurable weights, implemented in crates/libsy/src/algorithms/rand.rs. Accepts a list of model IDs, optional weights, and an optional seed for deterministic behavior.
  • algorithms.llm_classifier: Uses an LLM to classify and route requests based on content analysis.
  • algorithms.stage_router: Multi-stage routing that can chain different selection strategies.

Each factory accepts strategy-specific parameters, allowing fine-tuned control over the selection process without modifying the core Rust implementation in crates/libsy/src/algorithms.rs.

Summary

Embedding Switchyard's routing algorithms requires understanding the async iterator pattern and the separation between decision logic and execution:

  • Import routing factories from switchyard.libsy.algorithms to instantiate Rust-backed algorithms.
  • Implement an async LLM client that can handle both streaming and aggregated responses.
  • Iterate over algorithm.run_stream() and handle Step.CallModel by invoking your client and calling call.respond().
  • Process Step.Done to obtain the final selected_model_id and any cached response.
  • Reference the end-to-end example in examples/libsy.py for production-ready patterns.

Frequently Asked Questions

What Python version is required to use switchyard.libsy?

Switchyard requires Python 3.10 or later, as the integration examples utilize structural pattern matching (match/case syntax) introduced in PEP 634. The underlying PyO3 bindings in crates/switchyard-py/src/libsy_bindings.rs target the CPython stable ABI for compatibility with modern Python versions.

Can I use synchronous LLM clients with Switchyard's routing algorithms?

No, the Algorithm.run_stream() method is strictly async and returns an AsyncIterator[Step]. Your LLM client must provide async methods (using async def) to avoid blocking the event loop when handling Step.CallModel events. If your existing client is synchronous, wrap it with asyncio.to_thread() or similar executor patterns before passing responses to call.respond().

How does Step.CallModel differ from Step.Done in the routing lifecycle?

Step.CallModel represents an intermediate state where the algorithm has identified candidate models but needs you to execute the actual inference and return the results for evaluation. Step.Done represents the terminal state where the algorithm has finalized its decision and provides the RoutingOutcome containing the selected_model_id. A single routing pass may yield multiple CallModel steps before reaching Done.

Where is the actual routing logic implemented—Python or Rust?

The core routing logic is implemented in Rust within the libsy crate, specifically in files like crates/libsy/src/algorithms/rand.rs for random selection and crates/libsy/src/algorithms.rs for the trait definitions. The Python package switchyard.libsy provides thin wrappers via PyO3 that expose these Rust algorithms as Python objects, ensuring high-performance decision-making while allowing Python applications to control the I/O boundary.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →