Python Embedding Path vs Rust Algorithm::run_stream API in Switchyard: Key Differences

The Python embedding path provides a type-checked asynchronous iterator façade over the Rust core, where Algorithm::run_stream drives the actual routing logic and yields Step objects in both languages.

Switchyard’s routing engine bridges Rust performance with Python accessibility. Understanding the distinction between the high-level Python embedding path and the underlying Rust Algorithm::run_stream API is essential for developers integrating NVIDIA-NeMo/Switchyard into their inference pipelines.

Core Architecture and Implementation Locations

The routing logic resides in the Rust crate libsy, while Python users interact through a thin binding layer that mirrors the native interface.

Rust Core Implementation

In crates/libsy/src/core/algorithm.rs, the Algorithm::run_stream method defines the primary entry point. This method returns a StepStream implementing Stream<Item = Step>, where the Step enum at lines 352-362 consists of two variants: CallModel(ModelCall) and Done(RoutingOutcome). The Rust implementation handles model selection, routing logic, and response streaming without Python overhead.

Python Binding Layer

The Python interface is exposed through switchyard_rust/libsy.py, which defines type stubs for Algorithm, Step, and LlmResponse. The module switchyard/libsy/__init__.py re-exports these symbols for end-user consumption. These Python classes are thin wrappers around the Rust implementation, delegating all computational work to the native core via PyO3 bindings.

API Design and Type System Differences

While both APIs expose identical logical workflows, their type systems and calling conventions differ significantly.

Step and LlmResponse Type Definitions

Rust defines Step as a native enum in algorithm.rs with explicit memory layout:

pub enum Step {
    CallModel(ModelCall),
    Done(RoutingOutcome),
}

Python presents the same concept through the Step class defined in switchyard_rust/libsy.py (lines 42-52), providing Step.CallModel and Step.Done attributes for pattern matching. Similarly, LlmResponse exists as a Rust enum with Agg { response: Json } and Stream { stream: AsyncIterator } variants, while Python exposes LlmResponse.Agg and LlmResponse.Stream classes (lines 44-59).

Calling Patterns and Execution Flow

In Rust, developers interact directly with the asynchronous stream:

let algo = libsy::llm_task_classifier(config);
let mut stream = algo.run_stream(request, models);

while let Some(step) = stream.next().await {
    match step {
        Step::CallModel(call) => {
            // Execute model call and respond
            let response = execute_model(call.request).await;
            call.respond(LlmResponse::Agg(response));
        }
        Step::Done(outcome) => {
            println!("{:?}", outcome.response);
        }
    }
}

Python provides an identical asynchronous iteration interface through the binding layer:

from switchyard.libsy import Algorithm, Step, LlmResponse, algorithms

algo = algorithms.llm_task_classifier(config=config)

async for step in algo.run_stream(request={"messages": [...]}, models={"default": ["gpt4"]}):
    match step:
        case Step.CallModel(call):
            response = await some_http_call(call.request)
            call.respond(LlmResponse.Agg(response))
        case Step.Done(outcome):
            final = outcome.response
            print("Result:", final)

Both approaches yield CallModel steps for each required inference operation, followed by a terminal Done step containing the RoutingOutcome with the aggregated LlmResponse.

Error Handling Mechanisms

Rust returns errors as Result<StepStream, LibsyError>, allowing explicit error handling through pattern matching on Result types. The Python binding automatically converts these Rust errors into Python exceptions, raising LibsyError or ContextWindowExceededError directly when the underlying Rust call fails. This eliminates the need for manual error unwrapping in Python while preserving detailed error context from the core.

Performance Characteristics and Use Cases

The choice between APIs depends on your integration requirements and performance constraints.

Memory and Latency Considerations

The Python embedding path maintains objects as thin wrappers; the heavy lifting—model selection, routing logic, and response streaming—remains in Rust. Consequently, Python users experience latency comparable to native Rust usage, with minimal overhead from the binding layer. Pure Rust implementations avoid any Python GIL interaction, making the Algorithm::run_stream API preferable for library developers building high-throughput services or contributing core algorithm enhancements.

When to Use Each API

Use the Python embedding path when:

  • Integrating with existing Python frameworks like LiteLLM
  • Writing automation scripts or prototyping routing strategies (see examples/libsy.py)
  • Consuming the library as an end-user without modifying core logic

Use the Rust Algorithm::run_stream API when:

  • Implementing new routing algorithms in crates/libsy/src/algorithms/
  • Building performance-critical services that require zero Python overhead
  • Contributing to the crates/libsy library itself

Summary

  • The Rust core in crates/libsy/src/core/algorithm.rs implements Algorithm::run_stream as the canonical routing driver, returning a StepStream of native Step enums.
  • The Python binding in switchyard_rust/libsy.py provides type-checked stubs that re-export functionality through switchyard/libsy/__init__.py, maintaining identical semantics to the Rust API.
  • Both APIs yield CallModel steps for inference requests and a final Done step containing the RoutingOutcome, but Rust uses native enums while Python provides class-based pattern matching.
  • Error handling in Rust uses Result<StepStream, LibsyError> types, while Python raises LibsyError and ContextWindowExceededError exceptions converted from Rust errors.
  • Performance remains equivalent between paths because Python objects are thin wrappers; all computational work executes in the Rust core.

Frequently Asked Questions

What is the primary entry point for routing algorithms in Switchyard?

The primary entry point is the Algorithm::run_stream method defined in crates/libsy/src/core/algorithm.rs. This method drives the routing logic and returns an asynchronous stream of Step objects, available to both Rust consumers and Python users through the binding layer in switchyard/libsy/__init__.py.

How does error handling differ between Python and Rust in Switchyard?

Rust code returns Result<StepStream, LibsyError> types that developers must explicitly handle or unwrap using standard Rust patterns. In Python, the binding layer automatically converts these Rust errors into Python exceptions, specifically raising LibsyError or ContextWindowExceededError without requiring manual result unwrapping.

Can I use the Rust API directly without Python bindings?

Yes. The Rust API in crates/libsy operates independently of Python. Developers building pure-Rust applications or services can import libsy directly and use Algorithm::run_stream without any Python runtime or GIL constraints, as demonstrated in the crate's internal tests and algorithm implementations.

Where are the Step and LlmResponse types defined in the source code?

In Rust, these types are defined in crates/libsy/src/core/algorithm.rs at lines 352-362 for Step and lines 44-59 for LlmResponse. The Python type stubs appear in switchyard_rust/libsy.py at lines 42-52 for Step and lines 44-59 for LlmResponse, providing the interface that switchyard/libsy/__init__.py exposes to end users.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →