# Python Embedding Path vs Rust Algorithm::run_stream API in Switchyard: Key Differences

> Understand the key differences between Python embedding path (Step/LlmResponse) and Rust Algorithm::run_stream API in NVIDIA Switchyard. Learn how Python offers a type-checked facade over the Rust core.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**The Python embedding path provides a type-checked asynchronous iterator façade over the Rust core, where Algorithm::run_stream drives the actual routing logic and yields Step objects in both languages.**

Switchyard’s routing engine bridges Rust performance with Python accessibility. Understanding the distinction between the high-level Python embedding path and the underlying Rust Algorithm::run_stream API is essential for developers integrating NVIDIA-NeMo/Switchyard into their inference pipelines.

## Core Architecture and Implementation Locations

The routing logic resides in the Rust crate `libsy`, while Python users interact through a thin binding layer that mirrors the native interface.

### Rust Core Implementation

In [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), the `Algorithm::run_stream` method defines the primary entry point. This method returns a `StepStream` implementing `Stream<Item = Step>`, where the `Step` enum at lines 352-362 consists of two variants: `CallModel(ModelCall)` and `Done(RoutingOutcome)`. The Rust implementation handles model selection, routing logic, and response streaming without Python overhead.

### Python Binding Layer

The Python interface is exposed through [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), which defines type stubs for `Algorithm`, `Step`, and `LlmResponse`. The module [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py) re-exports these symbols for end-user consumption. These Python classes are thin wrappers around the Rust implementation, delegating all computational work to the native core via PyO3 bindings.

## API Design and Type System Differences

While both APIs expose identical logical workflows, their type systems and calling conventions differ significantly.

### Step and LlmResponse Type Definitions

Rust defines `Step` as a native enum in [`algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/algorithm.rs) with explicit memory layout:

```rust
pub enum Step {
    CallModel(ModelCall),
    Done(RoutingOutcome),
}

```

Python presents the same concept through the `Step` class defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) (lines 42-52), providing `Step.CallModel` and `Step.Done` attributes for pattern matching. Similarly, `LlmResponse` exists as a Rust enum with `Agg { response: Json }` and `Stream { stream: AsyncIterator }` variants, while Python exposes `LlmResponse.Agg` and `LlmResponse.Stream` classes (lines 44-59).

### Calling Patterns and Execution Flow

In Rust, developers interact directly with the asynchronous stream:

```rust
let algo = libsy::llm_task_classifier(config);
let mut stream = algo.run_stream(request, models);

while let Some(step) = stream.next().await {
    match step {
        Step::CallModel(call) => {
            // Execute model call and respond
            let response = execute_model(call.request).await;
            call.respond(LlmResponse::Agg(response));
        }
        Step::Done(outcome) => {
            println!("{:?}", outcome.response);
        }
    }
}

```

Python provides an identical asynchronous iteration interface through the binding layer:

```python
from switchyard.libsy import Algorithm, Step, LlmResponse, algorithms

algo = algorithms.llm_task_classifier(config=config)

async for step in algo.run_stream(request={"messages": [...]}, models={"default": ["gpt4"]}):
    match step:
        case Step.CallModel(call):
            response = await some_http_call(call.request)
            call.respond(LlmResponse.Agg(response))
        case Step.Done(outcome):
            final = outcome.response
            print("Result:", final)

```

Both approaches yield `CallModel` steps for each required inference operation, followed by a terminal `Done` step containing the `RoutingOutcome` with the aggregated `LlmResponse`.

### Error Handling Mechanisms

Rust returns errors as `Result<StepStream, LibsyError>`, allowing explicit error handling through pattern matching on `Result` types. The Python binding automatically converts these Rust errors into Python exceptions, raising `LibsyError` or `ContextWindowExceededError` directly when the underlying Rust call fails. This eliminates the need for manual error unwrapping in Python while preserving detailed error context from the core.

## Performance Characteristics and Use Cases

The choice between APIs depends on your integration requirements and performance constraints.

### Memory and Latency Considerations

The Python embedding path maintains objects as thin wrappers; the heavy lifting—model selection, routing logic, and response streaming—remains in Rust. Consequently, Python users experience latency comparable to native Rust usage, with minimal overhead from the binding layer. Pure Rust implementations avoid any Python GIL interaction, making the `Algorithm::run_stream` API preferable for library developers building high-throughput services or contributing core algorithm enhancements.

### When to Use Each API

Use the **Python embedding path** when:

- Integrating with existing Python frameworks like LiteLLM
- Writing automation scripts or prototyping routing strategies (see [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py))
- Consuming the library as an end-user without modifying core logic

Use the **Rust Algorithm::run_stream API** when:

- Implementing new routing algorithms in `crates/libsy/src/algorithms/`
- Building performance-critical services that require zero Python overhead
- Contributing to the `crates/libsy` library itself

## Summary

- The **Rust core** in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) implements `Algorithm::run_stream` as the canonical routing driver, returning a `StepStream` of native `Step` enums.
- The **Python binding** in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) provides type-checked stubs that re-export functionality through [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py), maintaining identical semantics to the Rust API.
- Both APIs yield `CallModel` steps for inference requests and a final `Done` step containing the `RoutingOutcome`, but Rust uses native enums while Python provides class-based pattern matching.
- Error handling in Rust uses `Result<StepStream, LibsyError>` types, while Python raises `LibsyError` and `ContextWindowExceededError` exceptions converted from Rust errors.
- Performance remains equivalent between paths because Python objects are thin wrappers; all computational work executes in the Rust core.

## Frequently Asked Questions

### What is the primary entry point for routing algorithms in Switchyard?

The primary entry point is the `Algorithm::run_stream` method defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs). This method drives the routing logic and returns an asynchronous stream of `Step` objects, available to both Rust consumers and Python users through the binding layer in [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py).

### How does error handling differ between Python and Rust in Switchyard?

Rust code returns `Result<StepStream, LibsyError>` types that developers must explicitly handle or unwrap using standard Rust patterns. In Python, the binding layer automatically converts these Rust errors into Python exceptions, specifically raising `LibsyError` or `ContextWindowExceededError` without requiring manual result unwrapping.

### Can I use the Rust API directly without Python bindings?

Yes. The Rust API in `crates/libsy` operates independently of Python. Developers building pure-Rust applications or services can import `libsy` directly and use `Algorithm::run_stream` without any Python runtime or GIL constraints, as demonstrated in the crate's internal tests and algorithm implementations.

### Where are the Step and LlmResponse types defined in the source code?

In Rust, these types are defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) at lines 352-362 for `Step` and lines 44-59 for `LlmResponse`. The Python type stubs appear in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) at lines 42-52 for `Step` and lines 44-59 for `LlmResponse`, providing the interface that [`switchyard/libsy/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/__init__.py) exposes to end users.