# Step::CallModel and Step::Done in Switchyard's Host-Side Orchestration Loop

> Understand Step::CallModel and Step::Done in Switchyard's host orchestration loop. Learn how these control messages decouple routing from I/O and manage LLM invocation and request termination.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-09-12

---

**Step::CallModel and Step::Done are the two primary control messages that decouple Switchyard's routing algorithm from I/O operations, instructing the host when to invoke an LLM and when to terminate the request stream.**

Switchyard is an open-source LLM routing engine developed by NVIDIA that separates routing logic from network I/O through a stream-based orchestration pattern. The host-side loop processes an asynchronous sequence of steps produced by the core algorithm, with **Step::CallModel** and **Step::Done** serving as the fundamental control primitives that coordinate model execution throughout the request lifecycle.

## The Step Enum as a Control Interface

The orchestration contract centers on the `Step` enum defined in [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs). This enum establishes a strict boundary between the pure Rust routing algorithm and the host runtime, ensuring that the algorithm can focus solely on decision-making while the runner handles all external side effects.

### Core Variants

The algorithm yields two primary variants during execution:

- **Step::CallModel** — Contains the selected `ModelId` and routing metadata, signaling the host to initiate an HTTP request to the chosen LLM backend.
- **Step::Done** — Indicates that the algorithm has reached a terminal state, whether through successful completion, error conditions, or token limit exhaustion.

## Step::CallModel: Dispatching to LLM Backends

When the routing algorithm determines which model should process a request, it yields `Step::CallModel`. The host implementation in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs) matches this variant to extract the target identifier and construct the appropriate HTTP request.

The host delegates the actual network call to [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs), which manages connection pooling, authentication headers, and response streaming. As tokens arrive from the LLM, the host forwards them to the caller while the algorithm remains suspended, awaiting the next decision point.

## Step::Done: Terminating the Orchestration Cycle

`Step::Done` serves as the terminal signal for the orchestration loop. Upon receiving this variant, the host performs cleanup operations: aggregating final metadata (such as the model selection chain and token counts), closing open response channels, and returning control to the caller.

This design ensures that resource finalization occurs predictably, regardless of whether the LLM connection closed naturally or was interrupted by a routing decision.

## Host-Side Orchestration Implementation

The orchestration loop resides in the consumption of `Algorithm::run_stream` from [`crates/libsy/src/core.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core.rs). The host asynchronously iterates over this stream, handling each step variant to coordinate I/O without blocking the routing logic.

### Processing the Step Stream

The following pattern illustrates how the runner consumes steps to bridge algorithm decisions with network operations:

```rust
// Conceptual implementation based on crates/switchyard-runner/src/runner.rs
use libsy::{Algorithm, Step};

async fn orchestrate_request(
    algorithm: &mut Algorithm,
    request: Request,
    llm_client: &LlmClient,
) -> Result<Response, Error> {
    let mut stream = algorithm.run_stream(request);
    
    while let Some(step) = stream.next().await {
        match step {
            Step::CallModel { model_id, .. } => {
                let response = llm_client
                    .call(model_id, request_body)
                    .await?;
                forward_tokens(response).await?;
            }
            Step::Done => {
                finalize_response();
                break;
            }
        }
    }
    
    Ok(build_final_response())
}

```

This architecture enables deterministic testing of the routing algorithm without network dependencies, as test harnesses can feed predefined step sequences instead of requiring live LLM endpoints.

## Summary

- **Step::CallModel** instructs the host to execute an HTTP request to the selected LLM backend, defined in [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs) and handled in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs).
- **Step::Done** signals request completion, allowing the host to clean up resources and return the final aggregated response.
- The orchestration loop in `Algorithm::run_stream` (from [`crates/libsy/src/core.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core.rs)) yields these steps to separate routing logic from I/O operations.
- This architecture enables the use of `libsy-llm-client` for actual network calls while keeping the algorithm pure and testable.

## Frequently Asked Questions

### What happens when the algorithm yields Step::CallModel?

The host runtime extracts the `ModelId` and associated routing metadata from the variant, then invokes the LLM client located in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs). The client establishes the HTTP connection, streams tokens back through the host, and maintains the connection until the model completes generation or the host receives a cancellation signal.

### How does Step::Done affect active streaming responses?

When the host encounters `Step::Done`, it immediately ceases processing the step stream, closes any open response channels, and triggers finalization logic to aggregate metadata such as the model selection chain and token counts. This ensures clean termination even if the underlying LLM connection remains technically active.

### Where are Step::CallModel and Step::Done defined in the Switchyard codebase?

Both variants are defined in the `Step` enum within [`crates/libsy/src/algorithms.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms.rs). The algorithm implementation in [`crates/libsy/src/core.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core.rs) instantiates these variants during the `run_stream` execution, while the runner in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs) consumes them to coordinate host-side I/O operations.

### Why does Switchyard use a step-based orchestration pattern instead of direct function calls?

The step-based design decouples the pure routing algorithm from asynchronous network I/O, allowing the algorithm to be unit tested with mock step sequences rather than requiring live LLM endpoints. This separation of concerns also enables the runner to handle connection management, retries, and streaming independently from the routing logic defined in the core library.