Understanding the Step Protocol in Switchyard: CallModel, Done, and the respond Contract

The Step protocol in Switchyard is a two-variant enum that bridges algorithmic routing logic with host-side model execution, where Step::CallModel pauses the algorithm to request a model call and Step::Done signals completion; the CallModel::respond contract requires a single, consuming call passing a Result<Response> to unblock the algorithm and either continue routing or trigger candidate fallback.

The NVIDIA-NeMo/Switchyard repository implements a model routing engine that separates pure routing algorithms from I/O-bound execution through a well-defined async protocol. This Step protocol in Switchyard, defined in the libsy crate, allows the Algorithm::run_stream method to yield control to the host whenever a model inference is required, then resume once the host fulfills the promise.

What Is the Step Protocol in Switchyard?

The Step protocol emerges from the implementation of Algorithm::run_stream in crates/libsy/src/core/algorithm.rs. Rather than blocking internally on network requests, the algorithm emits a stream of Step items, delegating all model invocations to the host environment through a pull-based async iterator. This design keeps the algorithm core agnostic of transport mechanisms while enabling sophisticated retry, fallback, and evaluation logic.

The Step enum is defined at lines 352-359 and serves as the sole communication channel between the routing algorithm and the driver:

pub enum Step {
    /// The routing algorithm needs a model to be called.
    CallModel(Box<CallModel>),
    /// The routing algorithm has finished.
    Done(Box<RoutingOutcome>),
}

Step Variants: CallModel and Done

The protocol operates through two distinct lifecycle phases:

  • Step::CallModel(Box): Emitted when the algorithm requires a model inference to proceed. This variant contains a boxed CallModel struct that carries the request payload, ordered candidate models, and a reply channel. The algorithm pauses at this step until the host invokes CallModel::respond, making this a synchronous barrier within the async stream.

  • Step::Done(Box): The terminal variant indicating the algorithm has finalized its routing decision. Once emitted, no further steps follow. The contained RoutingOutcome includes the selected model IDs, the original request, and optionally the final response text.

The CallModel Struct

The CallModel struct, defined at lines 95-114 in crates/libsy/src/core/algorithm.rs, encapsulates everything required to execute and respond to a model call:

pub struct CallModel {
    /// Name of the algorithm that emitted this call (useful for tracing).
    pub algorithm: String,
    /// The request that will be sent to the model. Its `model` field is stamped with the first candidate.
    pub request: Request,
    /// Ordered list of candidate `ModelId`s. The list is never empty.
    pub models: Vec<ModelId>,
    // Internal channel used to send the result back to the algorithm.
    reply: oneshot::Sender<Result<Response>>,
}

The struct provides the host with the algorithm name for observability, the specific request to send (targeting the first candidate in the list), and the ordered vector of fallback models. The private reply field holds a oneshot::Sender, enforcing the single-response contract through Rust's ownership system.

CallModel::respond Contract

The respond method, implemented at lines 116-124 in crates/libsy/src/core/algorithm.rs, fulfills the promise created by a Step::CallModel emission. The method signature and implementation reveal a strict contract:

impl CallModel {
    /// Fulfill the promise with the caller's model‑call result.
    /// Pass `Err(..)` to propagate a failed model call back to the algorithm.
    /// Consumes the promise: it can only be fulfilled once.
    pub fn respond(self, result: Result<Response>) -> Result<()> {
        self.reply
            .send(result)
            .map_err(|_| DriverError::ResponseDropped.into())
    }
}

The contract enforces four critical requirements:

  1. Single consumption: The method takes ownership of self, ensuring the promise can be fulfilled exactly once. Attempting to call respond twice results in a compile-time error due to move semantics.

  2. Result transmission: The parameter result: Result<Response> accepts either Ok(response) for successful model calls or Err(e) for model-level failures. Passing an error signals the algorithm to attempt the next candidate in the models vector, enabling automatic fallback.

  3. Infrastructure error handling: If the algorithm's receiver has been dropped (e.g., due to cancellation or timeout), respond returns DriverError::ResponseDropped. This indicates a fatal infrastructure error rather than a model error, as the algorithm can no longer receive the answer.

  4. Blocking semantics: The algorithm task remains paused at the CallModel step until respond is invoked. Once called, the oneshot channel transmits the result, unblocking the algorithm to continue processing the stream or emit Step::Done.

Practical Implementation Examples

Consuming Step::CallModel Manually

When implementing a custom host driver, you consume the step stream and handle Step::CallModel by executing the model call and fulfilling the promise:

use switchyard_libsy::{
    drive, Algorithm, CallModel, Step, RoutingOutcome,
};
use futures::stream::StreamExt;

async fn consume_steps<A: Algorithm>(
    alg: Arc<A>, 
    request: Request, 
    models: Arc<RuntimeModels>
) -> Result<RoutingOutcome> {
    // Define the serve closure that executes model calls
    let serve = |call: CallModel| async move {
        // Execute your LLM client logic here
        let response = my_llm_client
            .call(call.request, &call.models[0])
            .await;
        
        // Fulfill the promise - this unblocks the algorithm
        call.respond(response)
    };

    // Run the algorithm and drive the stream
    let outcome = drive(alg, request, models, serve).await?;
    println!("Routing completed: {:?}", outcome.selected_model_id());
    Ok(outcome)
}

In this pattern, the drive function from crates/libsy/src/core/algorithm.rs handles the stream polling, while your closure manages the actual HTTP or SDK request to the target LLM.

Using the Ready-Made Runner

For standard use cases, the switchyard-llm-client crate provides a high-level run function that implements the serve closure internally:

use switchyard_libsy::Algorithm;
use switchyard_libsy_llm_client::{run, ClientRouter};

async fn run_request<A: Algorithm>(
    algorithm: Arc<A>, 
    request: Request
) -> anyhow::Result<()> {
    let models = RuntimeModels::from(vec![/* available models */]);
    let router = ClientRouter::single(my_http_client);
    
    // run() handles Step::CallModel internally and returns on Step::Done
    let (selected_model, response) = run(
        algorithm, 
        router, 
        request, 
        Arc::new(models), 
        None
    ).await?;
    
    println!("Selected model: {}", selected_model);
    println!("Response: {}", response.text);
    Ok(())
}

This approach abstracts the Step protocol details while still relying on the same CallModel::respond contract internally.

Summary

  • The Step protocol in Switchyard consists of two variants emitted by Algorithm::run_stream: Step::CallModel requests model execution, while Step::Done indicates routing completion.
  • CallModel::respond consumes the struct and accepts a Result<Response>, unblocking the algorithm to continue or fall back to the next candidate.
  • The protocol enforces exactly-once semantics through Rust's ownership system, preventing double-response bugs at compile time.
  • DriverError::ResponseDropped signals fatal infrastructure failures when the algorithm receiver disappears before the host responds.
  • Source definitions reside in crates/libsy/src/core/algorithm.rs (lines 95-124 for CallModel, lines 352-359 for Step).

Frequently Asked Questions

What happens if CallModel::respond is called twice?

The respond method takes self by value, consuming the CallModel instance. Attempting to invoke it twice results in a Rust compile-time error because the ownership system prevents use-after-move, enforcing the exactly-once response contract statically.

How does the algorithm handle a failed model response?

When the host passes Err(e) to respond, the algorithm interprets this as a model-level failure (e.g., HTTP 5xx or timeout) and automatically retries with the next candidate in the CallModel.models vector. Only after exhausting all candidates does the algorithm surface the failure or emit Step::Done with an error state.

What is the difference between Step::CallModel and Step::Done?

Step::CallModel is an intermediate request for host-side execution that pauses the algorithm until respond is called, containing the request and candidate list. Step::Done is the terminal variant indicating no further processing is needed, containing the final RoutingOutcome with selected models and the response.

Where is the Step protocol defined in the Switchyard repository?

The protocol is defined in crates/libsy/src/core/algorithm.rs, specifically at lines 352-359 for the Step enum, lines 95-114 for the CallModel struct, and lines 116-124 for the respond method. The drive function consuming these steps appears in the same file, while crates/libsy-llm-client/src/run.rs provides the high-level implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →