How Switchyard's Driver::call_model Offloads Model Execution to the Host Harness
Switchyard's Driver::call_model offloads LLM inference to an external host harness by publishing a CallModel step through a promise-based channel, allowing the pure routing algorithm to await results without performing I/O directly.
Switchyard decouples routing algorithms from actual LLM inference through a clean separation of concerns implemented in the libsy crate. The Driver::call_model method serves as the bridge between algorithm logic and host-side execution, ensuring that routing decisions remain pure and testable while I/O operations happen externally according to the NVIDIA-NeMo/Switchyard source code.
The Architecture of Model Offloading
Separating Algorithm Logic from Inference
The routing algorithm runs inside the libsy crate and never performs network I/O directly. Instead, it creates a Driver instance that manages communication with the host harness through a one-slot mpsc channel. This design ensures that the routing logic remains deterministic and fully unit-testable without requiring mock HTTP servers.
Promise-Based Asynchronous Communication
The offloading mechanism relies on oneshot channels to create a promise pattern. When Driver::call_model is invoked, it constructs a CallModel struct containing the request, candidate models, and a oneshot::Sender. The algorithm awaits the corresponding receiver while the host fulfills the promise by calling CallModel::respond, creating a clean async boundary between the two components.
Step-by-Step Execution Flow
Initialize the Driver
The algorithm begins by instantiating a Driver paired with a step receiver. In crates/libsy/src/core/algorithm.rs, the constructor sets up the communication channel:
let (driver, step_rx) = Driver::new(self.name(), models);
The Driver::new function establishes a one-slot mpsc channel that the algorithm uses to emit steps, while the host receives these steps through step_rx.
Publish the CallModel Step
Inside crates/libsy/src/core/algorithm.rs at lines 260-292, the Driver::call_model method handles the offload:
let response = driver
.call_model(request.clone(), vec![target.clone()])
.await?;
The method performs four critical operations:
- Stamps the first candidate model onto the request
- Creates a
oneshotchannel (reply) for the host to send back results - Builds a
CallModelstruct containing the algorithm name, stamped request, candidate list, and reply sender - Publishes
Step::CallModel(Box::new(call))on itsstep_txchannel
The algorithm then awaits the reply receiver, pausing execution until the host fulfills the promise.
Consume Steps in the Host
The host runs libsy::drive, implemented in crates/libsy/src/core/algorithm.rs at lines 514-538. This function iterates over the algorithm's StepStream and delegates to a user-provided closure:
pub async fn drive<F, Fut>(..., serve: F) -> Result<RoutingOutcome>
where
F: Fn(CallModel) -> Fut,
Fut: Future<Output = Result<()>>,
{ … }
When the driver receives Step::CallModel(call), it invokes the serve closure, which performs the actual HTTP or RPC inference request to the LLM service.
Fulfill the Promise
After obtaining the LLM response, the host calls CallModel::respond to unblock the algorithm:
call.respond(Ok(response))?;
This sends the Result<Response> through the oneshot channel, allowing the awaiting call_model to return and the algorithm to continue processing or terminate with Step::Done.
Implementing the Host Harness
Custom Serve Function
A minimal host implementation provides a serve function that bridges Switchyard steps with real model endpoints. This function receives the CallModel, executes the inference, and fulfills the promise:
use libsy::{drive, CallModel, Result};
async fn serve(call: CallModel) -> Result<()> {
// Real inference request (e.g., HTTP to an LLM service)
let llm_resp = my_llm_client.call(call.request).await?;
// Fulfill the promise so the algorithm can continue
call.respond(Ok(llm_resp))
}
// Run the algorithm
let algo = std::sync::Arc::new(EchoAlgo);
let outcome = drive(algo, request, models, serve).await?;
Production-Ready Client
For production deployments, the switchyard-llm-client crate in crates/libsy-llm-client/src/run.rs provides a ready-made consumer that handles HTTP calls, retries, and automatic promise fulfillment without requiring custom serve implementations.
Concurrent Execution and Error Handling
Concurrent Hedging Pattern
Multiple call_model invocations can run concurrently without blocking the step channel. The algorithm can hedge by calling multiple models simultaneously and awaiting the first successful response:
let win = driver.call_model(req.clone(), vec!["fast".into()]);
let lose = driver.call_model(req, vec!["slow".into()]);
tokio::select! {
Ok(resp) = win => RoutingOutcome::answered("fast".into(), req, resp),
Ok(resp) = lose => RoutingOutcome::answered("slow".into(), req, resp),
}
Each call creates an independent oneshot channel, enabling parallel strategies while keeping the algorithm single-threaded and deterministic.
Error Propagation and Scope Handling
If the host drops the CallModel without responding, CallModel::respond returns DriverError::ResponseDropped. This propagates back to the algorithm as a step-level error, enabling graceful degradation or retry logic. Drivers also support scoping (Parent vs Subagent) to handle delegated sub-agents using the same offload flow.
Observability Integration
The call_model method is wrapped in a tracing::instrument span named libsy.llm_call and records latency, token usage, and outcome via observability::record_llm_call. This provides full visibility into offloaded requests without polluting the algorithm code with metrics calls.
Summary
- Driver::call_model publishes
CallModelsteps through anmpscchannel, never performing I/O directly inside the routing algorithm. - Oneshot channels create a promise pattern where the algorithm awaits results while the host fulfills requests through
CallModel::respond. - libsy::drive consumes the step stream in
crates/libsy/src/core/algorithm.rs(lines 514-538) and delegates to a user-providedservefunction for actual inference. - Concurrent calls are supported naturally through independent channels, enabling hedging and parallel routing strategies.
- DriverError::ResponseDropped handles host-side failures gracefully, while
tracingspans provide observability into the offload boundary.
Frequently Asked Questions
What file contains the Driver::call_model implementation?
The implementation resides in crates/libsy/src/core/algorithm.rs at lines 260-292. This code constructs the CallModel struct, manages the oneshot channel communication, and publishes steps to the driver's step_tx sender.
How does the host application receive and process CallModel steps?
The host calls libsy::drive, defined in crates/libsy/src/core/algorithm.rs at lines 514-538. This function accepts a serve closure that receives each CallModel instance. The host performs the actual HTTP or RPC request within this closure and calls CallModel::respond to return results through the oneshot channel.
Can multiple model calls run simultaneously from the same algorithm?
Yes. Each call_model invocation creates an independent oneshot channel pair, allowing the algorithm to issue concurrent requests and await them using tokio::select! or join! macros. The step channel remains unblocked because each call only publishes a lightweight CallModel struct without awaiting the response.
What happens if the host fails to respond or crashes?
If the host drops the CallModel without calling respond, the oneshot::Receiver closes. When the algorithm awaits the result, it receives DriverError::ResponseDropped, allowing the routing logic to handle the failure as a step-level error and potentially retry with alternative candidates or return a degraded response.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →