# How Python Applications Interact with Switchyard's Routing Algorithms Using switchyard-py

> Learn how Python applications interact with Switchyard's routing algorithms using the switchyard-py package. Explore PyO3 bindings for seamless integration and efficient routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Python applications interact with Switchyard's routing algorithms by importing factory functions from the `switchyard-py` package (distributed as `nemo-switchyard`), which exposes Rust-implemented algorithms through PyO3 bindings in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).**

Switchyard is an intelligent routing layer for LLM workloads developed by NVIDIA. The `switchyard-py` Python bindings allow developers to leverage high-performance routing algorithms—implemented in the Rust `libsy` crate—without writing any Rust code, enabling seamless integration into existing async Python applications.

## Architecture of the Python-Rust Bridge

Switchyard's Python integration relies on a three-tier architecture that isolates the compiled Rust logic from user-facing Python APIs.

- **`crates/libsy`** – The core Rust crate implementing routing algorithms (random selection, LLM-based classification, stage routers) and the streaming protocol.
- **[`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)** – A Python façade that lazily loads the compiled native library via `load_native()` and re-exports Rust symbols (`Algorithm`, `Step`, `LlmResponse`). It uses `__getattr__` to forward calls to the native library while providing type stubs for static analysis.
- **[`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)** – A convenience wrapper that imports factory functions (`random`, `llm_classifier`, `stage_router`) from `switchyard_rust.libsy` and exposes them in the public `switchyard.libsy.algorithms` namespace.

This design ensures that all heavy computational work occurs in optimized Rust code, while Python handles orchestration and I/O.

## Installing switchyard-py

Install the package via pip using its distribution name:

```bash
pip install nemo-switchyard

```

The package installs the compiled Rust extensions alongside the Python façade modules.

## Using Routing Algorithms in Python

### Importing the Algorithm Factories

Begin by importing the algorithms module, which provides factory functions for creating router instances:

```python
from switchyard.libsy import algorithms

```

This import path resolves through [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), which ultimately loads the Rust-backed implementations via [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).

### Creating a Routing Algorithm Instance

Instantiate a specific router by calling one of the factory functions. For example, create a **random router** that selects between targets `"fast"` and `"quality"` with weighted probabilities:

```python
algorithm = algorithms.random(
    ["fast", "quality"], 
    weights=[1, 3], 
    seed=42
)

```

The returned object is an `Algorithm` instance backed by the Rust implementation.

### Running the Algorithm Stream

Execute the routing logic by calling `run_stream()`, which returns an async iterator of `Step` objects:

```python
async for step in algorithm.run_stream(request):
    ...

```

The `request` parameter is a dictionary containing the LLM request payload (e.g., `{"model": "auto", "messages": [...]}`).

### Handling Step Types

The async iterator yields `Step` objects representing distinct phases of the routing lifecycle:

- **`Step.CallModel`** – Contains a `ModelCall` with the original request, a list of candidate models, and a `respond` callback. Your application must call the specified model and forward the response via `call.respond()`.
- **`Step.Done`** – Contains a `RoutingOutcome` exposing `selected_model_id`, `fallback_models`, and the final `response` (either an aggregate or stream).

## Complete Implementation Examples

### Random Router with Echo Client

The following example from [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) demonstrates a complete integration using a mock Echo client:

```python
#!/usr/bin/env python3
import asyncio
from collections.abc import AsyncIterator, Mapping

from switchyard.libsy import LlmResponse, Step, algorithms

class EchoClient:
    """Returns a fixed completion for any selected target."""
    async def call(self, request: Mapping[str, object], model: str) -> LlmResponse.Agg | LlmResponse.Stream:
        if request.get("stream"):
            async def events() -> AsyncIterator[Mapping[str, object]]:
                yield {"preservation": None, "normalized": [{"MessageStart": {"id": "echo", "model": model}}]}
                yield {"preservation": None, "normalized": [{"TextDelta": {"index": 0, "text": "Hello"}}]}
                yield {"preservation": None, "normalized": [{"MessageStop": {"reason": "end_turn"}}]}
            return LlmResponse.Stream(events())
        return LlmResponse.Agg(
            {
                "model": model,
                "outputs": [{"role": "assistant", "content": [{"type": "text", "text": "Hello"}]}],
            }
        )

async def main() -> None:
    request = {
        "model": "auto",
        "stream": True,
        "messages": [{"role": "user", "content": [{"type": "text", "text": "Hello"}]}],
    }
    client = EchoClient()
    # Build a random router that chooses between two targets.

    algorithm = algorithms.random(["fast", "quality"], weights=[1, 3], seed=42)

    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Forward the request to the selected model via the client.

                call.respond(await client.call(call.request, call.models[0]))
            case Step.Done(outcome):
                print("Decision:", outcome.selected_model_id)
                # The outcome may already contain a response; otherwise we ask the client again.

                response = outcome.response or await client.call(outcome.request, outcome.selected_model_id)
                match response:
                    case LlmResponse.Agg(agg):
                        print("Response:", agg)
                    case LlmResponse.Stream(stream):
                        async for ev in stream:
                            print("Response event:", ev)

if __name__ == "__main__":
    asyncio.run(main())

```

### LLM-Based Classifier

Configure a classifier router that uses an LLM to determine task capability:

```python
from switchyard.libsy import algorithms, LlmResponse, Step

config = algorithms.llm_classifier.capability(
    judge_target="judge",
    efficient_target="fast",
    capable_target="quality",
    config=algorithms.TaskClassifierConfig(
        base_threshold=0.7,
        session_affinity=False,
        recent_turn_window=10,
        max_output_tokens=1024,
        prompt="Classify the user request",
    ),
)

algorithm = algorithms.llm_classifier(config)

async for step in algorithm.run_stream(your_request):
    # Handle Step.CallModel and Step.Done as shown in the random router example

    pass

```

The `TaskClassifierConfig` class is defined in the type stubs within [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).

### Stage Router with Escalation

Implement a cascading router that escalates from an efficient model to a capable model based on confidence thresholds:

```python
algorithm = algorithms.stage_router(
    capable_target="quality",
    efficient_target="fast",
    picker="confidence",
    confidence_threshold=0.8,
    recent_window=20,
    escalation_note="Escalating due to low confidence",
    deescalation_note="De‑escalating after successful response",
)

```

This algorithm first attempts routing to the `efficient_target`, then escalates to `capable_target` if confidence falls below the specified threshold.

## Key Source Files and API Surface

| File | Purpose |
|------|---------|
| [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) | Python façade that lazily loads the compiled Rust `libsy` library and re-exports symbols (`Algorithm`, `Step`, `LlmResponse`). |
| [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) | Public factory module providing `random()`, `llm_classifier()`, and `stage_router()` functions. |
| [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) | End-to-end reference implementation showing algorithm initialization, mock client integration, and step handling. |
| [`docs/core_concepts.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/core_concepts.md) | Architectural documentation explaining the client-router-backend data flow. |

## Summary

- **switchyard-py** (pip package `nemo-switchyard`) exposes Switchyard's Rust routing algorithms to Python via PyO3 bindings in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).
- Applications interact with algorithms asynchronously using `run_stream()`, which yields `Step.CallModel` and `Step.Done` objects.
- The `switchyard.libsy.algorithms` module provides factory functions (`random`, `llm_classifier`, `stage_router`) that instantiate Rust-backed `Algorithm` objects.
- Users implement client logic to fulfill `CallModel` requests and feed responses back via the `respond()` callback.
- All heavy routing logic executes in the compiled Rust `libsy` crate, while Python manages async I/O and business logic.

## Frequently Asked Questions

### What is switchyard-py and how does it relate to the Switchyard project?

switchyard-py is the official Python client library for NVIDIA's Switchyard routing framework. It wraps the Rust `libsy` crate using PyO3 bindings, allowing Python applications to invoke Switchyard's routing algorithms without requiring Rust development expertise or toolchains.

### Do I need to write Rust code to use Switchyard's routing algorithms?

No. The `switchyard-py` package distributes pre-compiled Rust extensions. You interact with the algorithms entirely through Python imports from `switchyard.libsy` and `switchyard.libsy.algorithms`, as demonstrated in [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py).

### How do I handle streaming responses when using switchyard-py?

The `run_stream()` method returns an async iterator of `Step` objects. When you encounter `LlmResponse.Stream` (either in a `CallModel` step or the final `Done` outcome), iterate over its async generator to receive streaming chunks. For non-streaming use cases, handle `LlmResponse.Agg` to receive complete responses.

### Where are the routing algorithm implementations actually located?

The core algorithms reside in the Rust crate `crates/libsy`. The Python files ([`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) and [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py)) provide only thin façades and type stubs that forward calls to the compiled native library loaded at runtime.