# How Switchyard's Driver::call_model Offloads Model Execution to the Host Harness

> Learn how Switchyard's Driver::call_model offloads LLM inference to a host harness. Discover how it publishes steps via promise-based channels for efficient, I/O-free routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-09-12

---

**Switchyard's `Driver::call_model` offloads LLM inference to an external host harness by publishing a `CallModel` step through a promise-based channel, allowing the pure routing algorithm to await results without performing I/O directly.**

Switchyard decouples routing algorithms from actual LLM inference through a clean separation of concerns implemented in the `libsy` crate. The `Driver::call_model` method serves as the bridge between algorithm logic and host-side execution, ensuring that routing decisions remain pure and testable while I/O operations happen externally according to the NVIDIA-NeMo/Switchyard source code.

## The Architecture of Model Offloading

### Separating Algorithm Logic from Inference

The routing algorithm runs inside the `libsy` crate and never performs network I/O directly. Instead, it creates a **Driver** instance that manages communication with the host harness through a one-slot `mpsc` channel. This design ensures that the routing logic remains deterministic and fully unit-testable without requiring mock HTTP servers.

### Promise-Based Asynchronous Communication

The offloading mechanism relies on **oneshot channels** to create a promise pattern. When `Driver::call_model` is invoked, it constructs a `CallModel` struct containing the request, candidate models, and a `oneshot::Sender`. The algorithm awaits the corresponding receiver while the host fulfills the promise by calling `CallModel::respond`, creating a clean async boundary between the two components.

## Step-by-Step Execution Flow

### Initialize the Driver

The algorithm begins by instantiating a `Driver` paired with a step receiver. In [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), the constructor sets up the communication channel:

```rust
let (driver, step_rx) = Driver::new(self.name(), models);

```

The `Driver::new` function establishes a one-slot `mpsc` channel that the algorithm uses to emit steps, while the host receives these steps through `step_rx`.

### Publish the CallModel Step

Inside [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) at lines 260-292, the `Driver::call_model` method handles the offload:

```rust
let response = driver
    .call_model(request.clone(), vec![target.clone()])
    .await?;

```

The method performs four critical operations:
- Stamps the first candidate model onto the request
- Creates a `oneshot` channel (`reply`) for the host to send back results
- Builds a `CallModel` struct containing the algorithm name, stamped request, candidate list, and reply sender
- Publishes `Step::CallModel(Box::new(call))` on its `step_tx` channel

The algorithm then awaits the `reply` receiver, pausing execution until the host fulfills the promise.

### Consume Steps in the Host

The host runs `libsy::drive`, implemented in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) at lines 514-538. This function iterates over the algorithm's `StepStream` and delegates to a user-provided closure:

```rust
pub async fn drive<F, Fut>(..., serve: F) -> Result<RoutingOutcome>
where
    F: Fn(CallModel) -> Fut,
    Fut: Future<Output = Result<()>>,
{ … }

```

When the driver receives `Step::CallModel(call)`, it invokes the `serve` closure, which performs the actual HTTP or RPC inference request to the LLM service.

### Fulfill the Promise

After obtaining the LLM response, the host calls `CallModel::respond` to unblock the algorithm:

```rust
call.respond(Ok(response))?;

```

This sends the `Result<Response>` through the oneshot channel, allowing the awaiting `call_model` to return and the algorithm to continue processing or terminate with `Step::Done`.

## Implementing the Host Harness

### Custom Serve Function

A minimal host implementation provides a `serve` function that bridges Switchyard steps with real model endpoints. This function receives the `CallModel`, executes the inference, and fulfills the promise:

```rust
use libsy::{drive, CallModel, Result};

async fn serve(call: CallModel) -> Result<()> {
    // Real inference request (e.g., HTTP to an LLM service)
    let llm_resp = my_llm_client.call(call.request).await?;
    // Fulfill the promise so the algorithm can continue
    call.respond(Ok(llm_resp))
}

// Run the algorithm
let algo = std::sync::Arc::new(EchoAlgo);
let outcome = drive(algo, request, models, serve).await?;

```

### Production-Ready Client

For production deployments, the `switchyard-llm-client` crate in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) provides a ready-made consumer that handles HTTP calls, retries, and automatic promise fulfillment without requiring custom `serve` implementations.

## Concurrent Execution and Error Handling

### Concurrent Hedging Pattern

Multiple `call_model` invocations can run concurrently without blocking the step channel. The algorithm can hedge by calling multiple models simultaneously and awaiting the first successful response:

```rust
let win = driver.call_model(req.clone(), vec!["fast".into()]);
let lose = driver.call_model(req, vec!["slow".into()]);

tokio::select! {
    Ok(resp) = win => RoutingOutcome::answered("fast".into(), req, resp),
    Ok(resp) = lose => RoutingOutcome::answered("slow".into(), req, resp),
}

```

Each call creates an independent `oneshot` channel, enabling parallel strategies while keeping the algorithm single-threaded and deterministic.

### Error Propagation and Scope Handling

If the host drops the `CallModel` without responding, `CallModel::respond` returns **DriverError::ResponseDropped**. This propagates back to the algorithm as a step-level error, enabling graceful degradation or retry logic. Drivers also support scoping (**Parent** vs **Subagent**) to handle delegated sub-agents using the same offload flow.

### Observability Integration

The `call_model` method is wrapped in a `tracing::instrument` span named `libsy.llm_call` and records latency, token usage, and outcome via `observability::record_llm_call`. This provides full visibility into offloaded requests without polluting the algorithm code with metrics calls.

## Summary

- **Driver::call_model** publishes `CallModel` steps through an `mpsc` channel, never performing I/O directly inside the routing algorithm.
- **Oneshot channels** create a promise pattern where the algorithm awaits results while the host fulfills requests through `CallModel::respond`.
- **libsy::drive** consumes the step stream in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) (lines 514-538) and delegates to a user-provided `serve` function for actual inference.
- **Concurrent calls** are supported naturally through independent channels, enabling hedging and parallel routing strategies.
- **DriverError::ResponseDropped** handles host-side failures gracefully, while `tracing` spans provide observability into the offload boundary.

## Frequently Asked Questions

### What file contains the Driver::call_model implementation?

The implementation resides in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) at lines 260-292. This code constructs the `CallModel` struct, manages the oneshot channel communication, and publishes steps to the driver's `step_tx` sender.

### How does the host application receive and process CallModel steps?

The host calls `libsy::drive`, defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) at lines 514-538. This function accepts a `serve` closure that receives each `CallModel` instance. The host performs the actual HTTP or RPC request within this closure and calls `CallModel::respond` to return results through the oneshot channel.

### Can multiple model calls run simultaneously from the same algorithm?

Yes. Each `call_model` invocation creates an independent `oneshot` channel pair, allowing the algorithm to issue concurrent requests and await them using `tokio::select!` or `join!` macros. The step channel remains unblocked because each call only publishes a lightweight `CallModel` struct without awaiting the response.

### What happens if the host fails to respond or crashes?

If the host drops the `CallModel` without calling `respond`, the `oneshot::Receiver` closes. When the algorithm awaits the result, it receives `DriverError::ResponseDropped`, allowing the routing logic to handle the failure as a step-level error and potentially retry with alternative candidates or return a degraded response.