# Understanding the Step Protocol in Switchyard: CallModel, Done, and the respond Contract

> Explore the Switchyard Step protocol including CallModel and Done. Understand the CallModel respond contract for seamless routing and fallback.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**The Step protocol in Switchyard is a two-variant enum that bridges algorithmic routing logic with host-side model execution, where `Step::CallModel` pauses the algorithm to request a model call and `Step::Done` signals completion; the `CallModel::respond` contract requires a single, consuming call passing a `Result<Response>` to unblock the algorithm and either continue routing or trigger candidate fallback.**

The NVIDIA-NeMo/Switchyard repository implements a model routing engine that separates pure routing algorithms from I/O-bound execution through a well-defined async protocol. This **Step protocol in Switchyard**, defined in the `libsy` crate, allows the `Algorithm::run_stream` method to yield control to the host whenever a model inference is required, then resume once the host fulfills the promise.

## What Is the Step Protocol in Switchyard?

The Step protocol emerges from the implementation of `Algorithm::run_stream` in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs). Rather than blocking internally on network requests, the algorithm emits a stream of `Step` items, delegating all model invocations to the host environment through a pull-based async iterator. This design keeps the algorithm core agnostic of transport mechanisms while enabling sophisticated retry, fallback, and evaluation logic.

The `Step` enum is defined at lines 352-359 and serves as the sole communication channel between the routing algorithm and the driver:

```rust
pub enum Step {
    /// The routing algorithm needs a model to be called.
    CallModel(Box<CallModel>),
    /// The routing algorithm has finished.
    Done(Box<RoutingOutcome>),
}

```

### Step Variants: CallModel and Done

The protocol operates through two distinct lifecycle phases:

- **Step::CallModel(Box<CallModel>)**: Emitted when the algorithm requires a model inference to proceed. This variant contains a boxed `CallModel` struct that carries the request payload, ordered candidate models, and a reply channel. The algorithm pauses at this step until the host invokes `CallModel::respond`, making this a synchronous barrier within the async stream.

- **Step::Done(Box<RoutingOutcome>)**: The terminal variant indicating the algorithm has finalized its routing decision. Once emitted, no further steps follow. The contained `RoutingOutcome` includes the selected model IDs, the original request, and optionally the final response text.

## The CallModel Struct

The `CallModel` struct, defined at lines 95-114 in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), encapsulates everything required to execute and respond to a model call:

```rust
pub struct CallModel {
    /// Name of the algorithm that emitted this call (useful for tracing).
    pub algorithm: String,
    /// The request that will be sent to the model. Its `model` field is stamped with the first candidate.
    pub request: Request,
    /// Ordered list of candidate `ModelId`s. The list is never empty.
    pub models: Vec<ModelId>,
    // Internal channel used to send the result back to the algorithm.
    reply: oneshot::Sender<Result<Response>>,
}

```

The struct provides the host with the algorithm name for observability, the specific request to send (targeting the first candidate in the list), and the ordered vector of fallback models. The private `reply` field holds a `oneshot::Sender`, enforcing the single-response contract through Rust's ownership system.

## CallModel::respond Contract

The `respond` method, implemented at lines 116-124 in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), fulfills the promise created by a `Step::CallModel` emission. The method signature and implementation reveal a strict contract:

```rust
impl CallModel {
    /// Fulfill the promise with the caller's model‑call result.
    /// Pass `Err(..)` to propagate a failed model call back to the algorithm.
    /// Consumes the promise: it can only be fulfilled once.
    pub fn respond(self, result: Result<Response>) -> Result<()> {
        self.reply
            .send(result)
            .map_err(|_| DriverError::ResponseDropped.into())
    }
}

```

The contract enforces four critical requirements:

1. **Single consumption**: The method takes ownership of `self`, ensuring the promise can be fulfilled exactly once. Attempting to call `respond` twice results in a compile-time error due to move semantics.

2. **Result transmission**: The parameter `result: Result<Response>` accepts either `Ok(response)` for successful model calls or `Err(e)` for model-level failures. Passing an error signals the algorithm to attempt the next candidate in the `models` vector, enabling automatic fallback.

3. **Infrastructure error handling**: If the algorithm's receiver has been dropped (e.g., due to cancellation or timeout), `respond` returns `DriverError::ResponseDropped`. This indicates a fatal infrastructure error rather than a model error, as the algorithm can no longer receive the answer.

4. **Blocking semantics**: The algorithm task remains paused at the `CallModel` step until `respond` is invoked. Once called, the `oneshot` channel transmits the result, unblocking the algorithm to continue processing the stream or emit `Step::Done`.

## Practical Implementation Examples

### Consuming Step::CallModel Manually

When implementing a custom host driver, you consume the step stream and handle `Step::CallModel` by executing the model call and fulfilling the promise:

```rust
use switchyard_libsy::{
    drive, Algorithm, CallModel, Step, RoutingOutcome,
};
use futures::stream::StreamExt;

async fn consume_steps<A: Algorithm>(
    alg: Arc<A>, 
    request: Request, 
    models: Arc<RuntimeModels>
) -> Result<RoutingOutcome> {
    // Define the serve closure that executes model calls
    let serve = |call: CallModel| async move {
        // Execute your LLM client logic here
        let response = my_llm_client
            .call(call.request, &call.models[0])
            .await;
        
        // Fulfill the promise - this unblocks the algorithm
        call.respond(response)
    };

    // Run the algorithm and drive the stream
    let outcome = drive(alg, request, models, serve).await?;
    println!("Routing completed: {:?}", outcome.selected_model_id());
    Ok(outcome)
}

```

In this pattern, the `drive` function from [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) handles the stream polling, while your closure manages the actual HTTP or SDK request to the target LLM.

### Using the Ready-Made Runner

For standard use cases, the `switchyard-llm-client` crate provides a high-level `run` function that implements the `serve` closure internally:

```rust
use switchyard_libsy::Algorithm;
use switchyard_libsy_llm_client::{run, ClientRouter};

async fn run_request<A: Algorithm>(
    algorithm: Arc<A>, 
    request: Request
) -> anyhow::Result<()> {
    let models = RuntimeModels::from(vec![/* available models */]);
    let router = ClientRouter::single(my_http_client);
    
    // run() handles Step::CallModel internally and returns on Step::Done
    let (selected_model, response) = run(
        algorithm, 
        router, 
        request, 
        Arc::new(models), 
        None
    ).await?;
    
    println!("Selected model: {}", selected_model);
    println!("Response: {}", response.text);
    Ok(())
}

```

This approach abstracts the Step protocol details while still relying on the same `CallModel::respond` contract internally.

## Summary

- The **Step protocol** in Switchyard consists of two variants emitted by `Algorithm::run_stream`: `Step::CallModel` requests model execution, while `Step::Done` indicates routing completion.
- **`CallModel::respond`** consumes the struct and accepts a `Result<Response>`, unblocking the algorithm to continue or fall back to the next candidate.
- The protocol enforces **exactly-once semantics** through Rust's ownership system, preventing double-response bugs at compile time.
- **`DriverError::ResponseDropped`** signals fatal infrastructure failures when the algorithm receiver disappears before the host responds.
- Source definitions reside in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) (lines 95-124 for `CallModel`, lines 352-359 for `Step`).

## Frequently Asked Questions

### What happens if CallModel::respond is called twice?

The `respond` method takes `self` by value, consuming the `CallModel` instance. Attempting to invoke it twice results in a Rust compile-time error because the ownership system prevents use-after-move, enforcing the exactly-once response contract statically.

### How does the algorithm handle a failed model response?

When the host passes `Err(e)` to `respond`, the algorithm interprets this as a model-level failure (e.g., HTTP 5xx or timeout) and automatically retries with the next candidate in the `CallModel.models` vector. Only after exhausting all candidates does the algorithm surface the failure or emit `Step::Done` with an error state.

### What is the difference between Step::CallModel and Step::Done?

`Step::CallModel` is an intermediate request for host-side execution that pauses the algorithm until `respond` is called, containing the request and candidate list. `Step::Done` is the terminal variant indicating no further processing is needed, containing the final `RoutingOutcome` with selected models and the response.

### Where is the Step protocol defined in the Switchyard repository?

The protocol is defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), specifically at lines 352-359 for the `Step` enum, lines 95-114 for the `CallModel` struct, and lines 116-124 for the `respond` method. The `drive` function consuming these steps appears in the same file, while [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) provides the high-level implementation.