Understanding the Step Protocol in Switchyard: CallModel, Done, and the respond Contract
The Step protocol in Switchyard is a two-variant enum that bridges algorithmic routing logic with host-side model execution, where Step::CallModel pauses the algorithm to request a model call and Step::Done signals completion; the CallModel::respond contract requires a single, consuming call passing a Result<Response> to unblock the algorithm and either continue routing or trigger candidate fallback.
The NVIDIA-NeMo/Switchyard repository implements a model routing engine that separates pure routing algorithms from I/O-bound execution through a well-defined async protocol. This Step protocol in Switchyard, defined in the libsy crate, allows the Algorithm::run_stream method to yield control to the host whenever a model inference is required, then resume once the host fulfills the promise.
What Is the Step Protocol in Switchyard?
The Step protocol emerges from the implementation of Algorithm::run_stream in crates/libsy/src/core/algorithm.rs. Rather than blocking internally on network requests, the algorithm emits a stream of Step items, delegating all model invocations to the host environment through a pull-based async iterator. This design keeps the algorithm core agnostic of transport mechanisms while enabling sophisticated retry, fallback, and evaluation logic.
The Step enum is defined at lines 352-359 and serves as the sole communication channel between the routing algorithm and the driver:
pub enum Step {
/// The routing algorithm needs a model to be called.
CallModel(Box<CallModel>),
/// The routing algorithm has finished.
Done(Box<RoutingOutcome>),
}
Step Variants: CallModel and Done
The protocol operates through two distinct lifecycle phases:
-
Step::CallModel(Box): Emitted when the algorithm requires a model inference to proceed. This variant contains a boxed
CallModelstruct that carries the request payload, ordered candidate models, and a reply channel. The algorithm pauses at this step until the host invokesCallModel::respond, making this a synchronous barrier within the async stream. -
Step::Done(Box): The terminal variant indicating the algorithm has finalized its routing decision. Once emitted, no further steps follow. The contained
RoutingOutcomeincludes the selected model IDs, the original request, and optionally the final response text.
The CallModel Struct
The CallModel struct, defined at lines 95-114 in crates/libsy/src/core/algorithm.rs, encapsulates everything required to execute and respond to a model call:
pub struct CallModel {
/// Name of the algorithm that emitted this call (useful for tracing).
pub algorithm: String,
/// The request that will be sent to the model. Its `model` field is stamped with the first candidate.
pub request: Request,
/// Ordered list of candidate `ModelId`s. The list is never empty.
pub models: Vec<ModelId>,
// Internal channel used to send the result back to the algorithm.
reply: oneshot::Sender<Result<Response>>,
}
The struct provides the host with the algorithm name for observability, the specific request to send (targeting the first candidate in the list), and the ordered vector of fallback models. The private reply field holds a oneshot::Sender, enforcing the single-response contract through Rust's ownership system.
CallModel::respond Contract
The respond method, implemented at lines 116-124 in crates/libsy/src/core/algorithm.rs, fulfills the promise created by a Step::CallModel emission. The method signature and implementation reveal a strict contract:
impl CallModel {
/// Fulfill the promise with the caller's model‑call result.
/// Pass `Err(..)` to propagate a failed model call back to the algorithm.
/// Consumes the promise: it can only be fulfilled once.
pub fn respond(self, result: Result<Response>) -> Result<()> {
self.reply
.send(result)
.map_err(|_| DriverError::ResponseDropped.into())
}
}
The contract enforces four critical requirements:
-
Single consumption: The method takes ownership of
self, ensuring the promise can be fulfilled exactly once. Attempting to callrespondtwice results in a compile-time error due to move semantics. -
Result transmission: The parameter
result: Result<Response>accepts eitherOk(response)for successful model calls orErr(e)for model-level failures. Passing an error signals the algorithm to attempt the next candidate in themodelsvector, enabling automatic fallback. -
Infrastructure error handling: If the algorithm's receiver has been dropped (e.g., due to cancellation or timeout),
respondreturnsDriverError::ResponseDropped. This indicates a fatal infrastructure error rather than a model error, as the algorithm can no longer receive the answer. -
Blocking semantics: The algorithm task remains paused at the
CallModelstep untilrespondis invoked. Once called, theoneshotchannel transmits the result, unblocking the algorithm to continue processing the stream or emitStep::Done.
Practical Implementation Examples
Consuming Step::CallModel Manually
When implementing a custom host driver, you consume the step stream and handle Step::CallModel by executing the model call and fulfilling the promise:
use switchyard_libsy::{
drive, Algorithm, CallModel, Step, RoutingOutcome,
};
use futures::stream::StreamExt;
async fn consume_steps<A: Algorithm>(
alg: Arc<A>,
request: Request,
models: Arc<RuntimeModels>
) -> Result<RoutingOutcome> {
// Define the serve closure that executes model calls
let serve = |call: CallModel| async move {
// Execute your LLM client logic here
let response = my_llm_client
.call(call.request, &call.models[0])
.await;
// Fulfill the promise - this unblocks the algorithm
call.respond(response)
};
// Run the algorithm and drive the stream
let outcome = drive(alg, request, models, serve).await?;
println!("Routing completed: {:?}", outcome.selected_model_id());
Ok(outcome)
}
In this pattern, the drive function from crates/libsy/src/core/algorithm.rs handles the stream polling, while your closure manages the actual HTTP or SDK request to the target LLM.
Using the Ready-Made Runner
For standard use cases, the switchyard-llm-client crate provides a high-level run function that implements the serve closure internally:
use switchyard_libsy::Algorithm;
use switchyard_libsy_llm_client::{run, ClientRouter};
async fn run_request<A: Algorithm>(
algorithm: Arc<A>,
request: Request
) -> anyhow::Result<()> {
let models = RuntimeModels::from(vec![/* available models */]);
let router = ClientRouter::single(my_http_client);
// run() handles Step::CallModel internally and returns on Step::Done
let (selected_model, response) = run(
algorithm,
router,
request,
Arc::new(models),
None
).await?;
println!("Selected model: {}", selected_model);
println!("Response: {}", response.text);
Ok(())
}
This approach abstracts the Step protocol details while still relying on the same CallModel::respond contract internally.
Summary
- The Step protocol in Switchyard consists of two variants emitted by
Algorithm::run_stream:Step::CallModelrequests model execution, whileStep::Doneindicates routing completion. CallModel::respondconsumes the struct and accepts aResult<Response>, unblocking the algorithm to continue or fall back to the next candidate.- The protocol enforces exactly-once semantics through Rust's ownership system, preventing double-response bugs at compile time.
DriverError::ResponseDroppedsignals fatal infrastructure failures when the algorithm receiver disappears before the host responds.- Source definitions reside in
crates/libsy/src/core/algorithm.rs(lines 95-124 forCallModel, lines 352-359 forStep).
Frequently Asked Questions
What happens if CallModel::respond is called twice?
The respond method takes self by value, consuming the CallModel instance. Attempting to invoke it twice results in a Rust compile-time error because the ownership system prevents use-after-move, enforcing the exactly-once response contract statically.
How does the algorithm handle a failed model response?
When the host passes Err(e) to respond, the algorithm interprets this as a model-level failure (e.g., HTTP 5xx or timeout) and automatically retries with the next candidate in the CallModel.models vector. Only after exhausting all candidates does the algorithm surface the failure or emit Step::Done with an error state.
What is the difference between Step::CallModel and Step::Done?
Step::CallModel is an intermediate request for host-side execution that pauses the algorithm until respond is called, containing the request and candidate list. Step::Done is the terminal variant indicating no further processing is needed, containing the final RoutingOutcome with selected models and the response.
Where is the Step protocol defined in the Switchyard repository?
The protocol is defined in crates/libsy/src/core/algorithm.rs, specifically at lines 352-359 for the Step enum, lines 95-114 for the CallModel struct, and lines 116-124 for the respond method. The drive function consuming these steps appears in the same file, while crates/libsy-llm-client/src/run.rs provides the high-level implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →