Provider-Neutral Request and Response Types in Switchyard: The Core Abstraction Layer
Switchyard defines provider-neutral request and response types through the Request and Response envelope structs in the switchyard-protocol crate, which wrap normalized LlmRequest and LlmResponse payloads to abstract away provider-specific API differences.
The NVIDIA-NeMo/Switchyard repository implements a unified routing layer for LLM traffic across diverse providers like OpenAI, Anthropic, and NVIDIA. At the heart of this system lies a provider-neutral data model that enables routing algorithms to operate without knowledge of upstream API specifics. These normalized types serve as the exclusive interface between the translation layer, routing logic, and language bindings.
The Core Envelope Architecture
Switchyard encapsulates all LLM traffic within two primary envelope structures defined in crates/protocol/src/envelope.rs. These envelopes act as containers that wrap normalized payloads while preserving optional raw provider data and correlation metadata.
Request Envelope Structure
The Request struct represents the standardized input that routing algorithms receive. According to the source in crates/protocol/src/envelope.rs, this struct holds the normalized request alongside optional host-owned data:
/// A request an algorithm routes: the normalized [`LlmRequest`] plus optional
/// host‑owned raw data and correlation [`Metadata`].
#[derive(Clone, Default)]
pub struct Request {
/// The normalized request an algorithm routes.
pub llm_request: LlmRequest,
/// An optional whole request body retained by the host.
pub raw_request: Option<serde_json::Value>,
/// Correlation metadata carried through the request.
pub metadata: Option<Metadata>,
}
The raw_request field retains the original provider-specific payload, while llm_request contains the normalized representation used for routing decisions.
Response Envelope Structure
The Response struct mirrors the request pattern, wrapping the normalized LlmResponse with associated metadata:
/// A response an algorithm returns: the [`LlmResponse`] (streamed or aggregate)
/// plus optional correlation [`Metadata`].
pub struct Response {
/// The neutral model response — a chunk stream or the buffered aggregate.
pub llm_response: LlmResponse,
/// Correlation metadata carried through the response.
pub metadata: Option<Metadata>,
}
The implementation in crates/protocol/src/envelope.rs includes methods like served_model() and set_served_model() (lines 73–86) for tracking which model actually served the request during the round-trip.
Normalized Payload Types
The actual semantic content travels within the llm_request and llm_response fields, defined in crates/protocol/src/llm.rs. These types strip away provider-specific formatting while preserving essential LLM parameters.
LlmRequest Structure
The LlmRequest struct represents the normalized request sent to provider-agnostic routing algorithms:
/// The normalized request sent to a provider‑agnostic routing algorithm.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct LlmRequest {
/// The model identifier (optional – filled in by the routing driver).
pub model: Option<String>,
/// The conversation history as a list of messages.
pub messages: Vec<Message>,
/// Optional instruction block that is kept separate from the normal messages.
pub instruction: Option<InstructionBlock>,
/// Sampling, usage, and provider‑specific extensions.
pub sampling_params: Option<SamplingParams>,
pub usage: Option<Usage>,
pub extensions: Option<ProviderExtensions>,
}
This structure supports conversation history through the messages vector, separate system instructions via instruction, and provider-specific extensions through the extensions field.
LlmResponse Variants
The LlmResponse enum handles both aggregate and streaming response patterns:
/// The neutral response that an algorithm returns to the caller.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum LlmResponse {
/// Aggregate (fully buffered) response.
Agg(AggLlmResponse),
/// Streaming response, where the payload is an async iterator.
Stream(LlmResponseStreamEvent),
}
The Agg variant contains complete buffered responses, while Stream represents async iterators for token-by-token delivery. This dual-variant approach allows routing algorithms to handle both response modes through a single type.
Practical Implementation Examples
Creating a provider-neutral request in Rust involves constructing the envelope with the normalized payload:
use switchyard_protocol::{
envelope::Request,
llm::{LlmRequest, Message, Role},
metadata::Metadata,
};
let req = Request {
llm_request: LlmRequest {
model: Some("gpt-4".into()),
messages: vec![
Message::text(Role::User, "Explain quantum tunneling."),
],
..Default::default()
},
raw_request: None,
metadata: Some(Metadata::default()),
};
The Python bindings expose these same types through switchyard.libsy, allowing Python-based routing algorithms to interact with the neutral model:
from switchyard.libsy import LlmResponse, Step, algorithms
# Assume `call` is a `Step` produced by a routing algorithm.
call.respond(LlmResponse.Agg({"content": "Quantum tunneling …"}))
# Later, when handling the response:
match call.response.llm_request:
case LlmResponse.Agg(agg):
print("Full answer:", agg["content"])
case LlmResponse.Stream(stream):
for event in stream:
print("Streamed chunk:", event)
Round-tripping model identifiers demonstrates how the envelope preserves routing metadata:
let mut resp = Response {
llm_response: LlmResponse::Agg(text_response(None, "answer")),
metadata: None,
};
assert_eq!(resp.served_model(), None);
resp.set_served_model(&ModelId::from("first"));
assert_eq!(resp.served_model().map(ModelId::as_str), Some("first"));
Integration Points
The provider-neutral request and response types serve as the exclusive public interface for all Switchyard components. Routing algorithms in switchyard.libsy.algorithms receive Request objects and return Response objects without accessing provider-specific formats. The translation layer converts OpenAI, Anthropic, or NVIDIA API formats to these neutral structures before routing decisions occur, then converts back when sending responses to clients.
Summary
- Envelope Pattern: The
RequestandResponsestructs incrates/protocol/src/envelope.rswrap normalized payloads with optional raw data and metadata. - Normalized Types:
LlmRequestandLlmResponseincrates/protocol/src/llm.rsprovide a unified interface for LLM parameters and output formats. - Dual Response Modes:
LlmResponsesupports both aggregate (Agg) and streaming (Stream) variants through a single enum. - Language Bindings: Python access via
switchyard.libsyexposes the same Rust-backed types for cross-language algorithm development. - Metadata Preservation: Envelopes carry correlation metadata and model identifiers through the entire request lifecycle.
Frequently Asked Questions
What is the difference between the envelope and the normalized payload?
The envelope (Request/Response) is a container struct that holds the normalized payload plus additional routing metadata. The normalized payload (LlmRequest/LlmResponse) contains the actual LLM-specific data like messages, model identifiers, and response content. The envelope pattern allows Switchyard to attach correlation metadata and preserve raw provider data without polluting the normalized model.
How does Switchyard handle streaming responses in the neutral type system?
Switchyard represents streaming responses through the LlmResponse::Stream variant, which wraps an async iterator of LlmResponseStreamEvent. Routing algorithms can pattern-match on the response enum to handle either fully-buffered aggregate responses or token streams uniformly, without knowing whether the underlying provider supports server-sent events or WebSocket streaming.
Can provider-specific metadata be preserved when using these neutral types?
Yes. The Request struct includes a raw_request field of type Option<serde_json::Value> that retains the complete original provider payload. Additionally, both Request and Response carry an optional metadata field for correlation data. This design ensures no information is lost during normalization while still providing a clean interface for routing logic.
Where are these types defined in the Switchyard codebase?
The core provider-neutral types reside in the switchyard-protocol crate. Specifically, crates/protocol/src/envelope.rs defines the Request and Response envelopes, while crates/protocol/src/llm.rs contains the LlmRequest struct and LlmResponse enum. Python bindings exposing these types are located in switchyard/libsy/__init__.py with the low-level implementation in switchyard_rust/libsy.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →