Python Embedding Path vs Rust Algorithm::run_stream API in Switchyard: Key Differences
The Python embedding path provides a type-checked asynchronous iterator façade over the Rust core, where Algorithm::run_stream drives the actual routing logic and yields Step objects in both languages.
Switchyard’s routing engine bridges Rust performance with Python accessibility. Understanding the distinction between the high-level Python embedding path and the underlying Rust Algorithm::run_stream API is essential for developers integrating NVIDIA-NeMo/Switchyard into their inference pipelines.
Core Architecture and Implementation Locations
The routing logic resides in the Rust crate libsy, while Python users interact through a thin binding layer that mirrors the native interface.
Rust Core Implementation
In crates/libsy/src/core/algorithm.rs, the Algorithm::run_stream method defines the primary entry point. This method returns a StepStream implementing Stream<Item = Step>, where the Step enum at lines 352-362 consists of two variants: CallModel(ModelCall) and Done(RoutingOutcome). The Rust implementation handles model selection, routing logic, and response streaming without Python overhead.
Python Binding Layer
The Python interface is exposed through switchyard_rust/libsy.py, which defines type stubs for Algorithm, Step, and LlmResponse. The module switchyard/libsy/__init__.py re-exports these symbols for end-user consumption. These Python classes are thin wrappers around the Rust implementation, delegating all computational work to the native core via PyO3 bindings.
API Design and Type System Differences
While both APIs expose identical logical workflows, their type systems and calling conventions differ significantly.
Step and LlmResponse Type Definitions
Rust defines Step as a native enum in algorithm.rs with explicit memory layout:
pub enum Step {
CallModel(ModelCall),
Done(RoutingOutcome),
}
Python presents the same concept through the Step class defined in switchyard_rust/libsy.py (lines 42-52), providing Step.CallModel and Step.Done attributes for pattern matching. Similarly, LlmResponse exists as a Rust enum with Agg { response: Json } and Stream { stream: AsyncIterator } variants, while Python exposes LlmResponse.Agg and LlmResponse.Stream classes (lines 44-59).
Calling Patterns and Execution Flow
In Rust, developers interact directly with the asynchronous stream:
let algo = libsy::llm_task_classifier(config);
let mut stream = algo.run_stream(request, models);
while let Some(step) = stream.next().await {
match step {
Step::CallModel(call) => {
// Execute model call and respond
let response = execute_model(call.request).await;
call.respond(LlmResponse::Agg(response));
}
Step::Done(outcome) => {
println!("{:?}", outcome.response);
}
}
}
Python provides an identical asynchronous iteration interface through the binding layer:
from switchyard.libsy import Algorithm, Step, LlmResponse, algorithms
algo = algorithms.llm_task_classifier(config=config)
async for step in algo.run_stream(request={"messages": [...]}, models={"default": ["gpt4"]}):
match step:
case Step.CallModel(call):
response = await some_http_call(call.request)
call.respond(LlmResponse.Agg(response))
case Step.Done(outcome):
final = outcome.response
print("Result:", final)
Both approaches yield CallModel steps for each required inference operation, followed by a terminal Done step containing the RoutingOutcome with the aggregated LlmResponse.
Error Handling Mechanisms
Rust returns errors as Result<StepStream, LibsyError>, allowing explicit error handling through pattern matching on Result types. The Python binding automatically converts these Rust errors into Python exceptions, raising LibsyError or ContextWindowExceededError directly when the underlying Rust call fails. This eliminates the need for manual error unwrapping in Python while preserving detailed error context from the core.
Performance Characteristics and Use Cases
The choice between APIs depends on your integration requirements and performance constraints.
Memory and Latency Considerations
The Python embedding path maintains objects as thin wrappers; the heavy lifting—model selection, routing logic, and response streaming—remains in Rust. Consequently, Python users experience latency comparable to native Rust usage, with minimal overhead from the binding layer. Pure Rust implementations avoid any Python GIL interaction, making the Algorithm::run_stream API preferable for library developers building high-throughput services or contributing core algorithm enhancements.
When to Use Each API
Use the Python embedding path when:
- Integrating with existing Python frameworks like LiteLLM
- Writing automation scripts or prototyping routing strategies (see
examples/libsy.py) - Consuming the library as an end-user without modifying core logic
Use the Rust Algorithm::run_stream API when:
- Implementing new routing algorithms in
crates/libsy/src/algorithms/ - Building performance-critical services that require zero Python overhead
- Contributing to the
crates/libsylibrary itself
Summary
- The Rust core in
crates/libsy/src/core/algorithm.rsimplementsAlgorithm::run_streamas the canonical routing driver, returning aStepStreamof nativeStepenums. - The Python binding in
switchyard_rust/libsy.pyprovides type-checked stubs that re-export functionality throughswitchyard/libsy/__init__.py, maintaining identical semantics to the Rust API. - Both APIs yield
CallModelsteps for inference requests and a finalDonestep containing theRoutingOutcome, but Rust uses native enums while Python provides class-based pattern matching. - Error handling in Rust uses
Result<StepStream, LibsyError>types, while Python raisesLibsyErrorandContextWindowExceededErrorexceptions converted from Rust errors. - Performance remains equivalent between paths because Python objects are thin wrappers; all computational work executes in the Rust core.
Frequently Asked Questions
What is the primary entry point for routing algorithms in Switchyard?
The primary entry point is the Algorithm::run_stream method defined in crates/libsy/src/core/algorithm.rs. This method drives the routing logic and returns an asynchronous stream of Step objects, available to both Rust consumers and Python users through the binding layer in switchyard/libsy/__init__.py.
How does error handling differ between Python and Rust in Switchyard?
Rust code returns Result<StepStream, LibsyError> types that developers must explicitly handle or unwrap using standard Rust patterns. In Python, the binding layer automatically converts these Rust errors into Python exceptions, specifically raising LibsyError or ContextWindowExceededError without requiring manual result unwrapping.
Can I use the Rust API directly without Python bindings?
Yes. The Rust API in crates/libsy operates independently of Python. Developers building pure-Rust applications or services can import libsy directly and use Algorithm::run_stream without any Python runtime or GIL constraints, as demonstrated in the crate's internal tests and algorithm implementations.
Where are the Step and LlmResponse types defined in the source code?
In Rust, these types are defined in crates/libsy/src/core/algorithm.rs at lines 352-362 for Step and lines 44-59 for LlmResponse. The Python type stubs appear in switchyard_rust/libsy.py at lines 42-52 for Step and lines 44-59 for LlmResponse, providing the interface that switchyard/libsy/__init__.py exposes to end users.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →