How the Switchyard Protocol Crate Defines Provider-Neutral Request and Response Types
The protocol crate defines pure-Rust structs LlmRequest and AggLlmResponse that normalize LLM interactions from any provider into a provider-neutral format, enabling unified routing and lossless round-tripping through fields like extensions and preservation.
The NVIDIA-NeMo/Switchyard inference router relies on the protocol crate to establish a canonical representation of LLM traffic. By mapping every OpenAI, Anthropic, or NVIDIA native payload into these normalized types, Switchyard performs cross-provider routing, A/B testing, and logging without maintaining bespoke logic for each backend.
Core Request Type: LlmRequest
The primary request structure is defined in [crates/protocol/src/llm.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs) at line 103. The LlmRequest struct captures every concept common to modern LLM APIs while remaining agnostic to any specific provider's JSON schema.
Key Fields and Provider-Neutral Semantics
The struct contains eleven top-level fields that cover the full spectrum of inference parameters:
- model: An optional model identifier that routing layers may rewrite before forwarding to a specific backend.
- instructions: System or developer blocks separated from conversational turns, supporting providers that distinguish system-level prompts.
- messages: An ordered vector of
Messagestructs representing the conversation history, each containing aRoleand vector ofContentBlock. - tools: A list of tool definitions following a generic schema that mirrors OpenAI's function format but works across providers.
- tool_choice: A normalized policy enum (auto, required, none, or specific tool) that unifies various provider-specific tool selection fields.
- sampling: Contains
temperature,top_p, andtop_kparameters as a unifiedSamplingParamsstruct. - output: Specifies token budgets through
max_output_tokensand optionalresponse_formatconstraints. - reasoning: Captures provider-specific reasoning controls like
effortlevels or raw reasoning flags in aReasoningParamsstruct. - stream: A boolean flag that unifies streaming requests across all backends.
- extensions: A
ProviderExtensionsfield holding raw provider-specific fields that lack first-class representation, ensuring lossless round-tripping. - preservation:
PreservationMetadatathat stores exact request payloads keyed by source format, enabling same-format replay without re-encoding.
Supporting Types for Content and Roles
The request relies on several supporting enums defined in the same file:
Role: Normalizes actor identities intoSystem,Developer,User,Assistant, andToolvariants.ContentBlock: Represents message content through variants includingText,Reasoning,Image,Audio,Video,File,ToolCall,ToolResult,Refusal, and a genericUnknowncatch-all.Message: Combines aRolewith a vector ofContentBlockelements to represent a single conversation turn.
Additional supporting structs include ToolDefinition, ToolChoice, SamplingParams, OutputParams, ReasoningParams, ProviderExtensions, and PreservationMetadata.
Core Response Type: AggLlmResponse
The response counterpart is defined at line 38 of [crates/protocol/src/llm.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs). The AggLlmResponse struct provides a normalized view of LLM outputs that preserves provider-specific metadata while presenting a unified interface.
Response Structure and Usage Tracking
The response type contains six primary fields:
- id: An optional identifier returned by the provider, kept verbatim for traceability.
- model: The model name reported by the backend, which may differ from the request's model field after routing decisions.
- outputs: An ordered list of
ResponseOutputstructs, each containing a role, vector ofContentBlock, and optionalStopReason. - usage: A normalized
Usagestruct containinginput_tokens,output_tokens,total_tokens, cache details viaInputCacheUsage, and reasoning token counts. - extensions: Provider-specific response fields without first-class equivalents, stored in
ProviderExtensionsfor lossless round-tripping. - preservation: Exact response payloads keyed by source format, mirroring the request preservation mechanism.
Stop Reasons and Output Handling
Each ResponseOutput within the outputs vector includes a stop_reason field using the StopReason enum. This enum normalizes termination conditions across providers:
EndTurn: Natural completion of the response.MaxTokens: Generation halted due to token budget exhaustion.ToolUse: The model elected to invoke a tool.ContentFilter: Output blocked by safety filters.Error: Provider-reported error conditions.Unknown: Unrecognized stop conditions from novel providers.
Architecture Benefits of Provider-Neutral Types
Because every Switchyard backend—whether a Rust server, Python façade, or external provider client—first translates native payloads into these canonical types, the system gains several architectural advantages:
- Unified Routing: The router evaluates
model,messages, andsamplingparameters without parsing provider-specific JSON variants. - Lossless Transformation: The
extensionsandpreservationfields guarantee that unknown provider fields survive round-trips, enabling "same-format" replay scenarios where the exact original payload must be reconstructed. - Cross-Provider Logging: A single logging implementation can record
AggLlmResponsestructs with normalizedUsagemetrics regardless of whether the backend was OpenAI, Anthropic, or a local vLLM instance.
Working with Protocol Types in Code
All structs are re-exported from the crate root in [crates/protocol/src/lib.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), making them accessible as protocol::LlmRequest and protocol::AggLlmResponse.
Rust Examples
Constructing a request manually:
use switchyard_protocol::{LlmRequest, Message, Role, ContentBlock};
let req = LlmRequest {
model: Some("mixtral-8x7b".into()),
messages: vec![
Message::text(Role::User, "Write a Rust hello world program."),
],
..Default::default()
};
// Serialize to JSON for transport
let json = serde_json::to_string(&req).unwrap();
Handling a normalized response:
use switchyard_protocol::{AggLlmResponse, StopReason};
fn handle_response(resp: AggLlmResponse) {
if let Some(first) = resp.first_output() {
if let Some(StopReason::ToolUse) = first.stop_reason {
println!("Model requested a tool call");
}
}
}
Python Bindings
When using Switchyard's Python interface, the same types are available through generated bindings:
from switchyard.protocol import LlmRequest, Role, Message
# Build a simple single-turn request
request = LlmRequest.text_request(
model="gpt-4o-mini",
prompt="Explain the waterfall model."
)
print(request.model) # → "gpt-4o-mini"
print(request.messages[0].role) # → Role.User
Summary
- The
protocolcrate definesLlmRequest(line 103) andAggLlmResponse(line 38) incrates/protocol/src/llm.rsas the canonical provider-neutral types for all Switchyard traffic. LlmRequestunifies model selection, conversation history, tool definitions, sampling parameters, and reasoning controls while preserving unknown fields viaextensionsandpreservation.AggLlmResponsenormalizes response identifiers, outputs, token usage, and stop reasons across all LLM providers.- Supporting types like
ContentBlock,Role, andStopReasonabstract provider-specific representations into generic enums. - These types enable unified routing, logging, and lossless round-tripping regardless of the underlying LLM backend.
Frequently Asked Questions
How does the protocol crate handle provider-specific fields that don't map to standard LLM parameters?
The LlmRequest and AggLlmResponse structs include an extensions field of type ProviderExtensions that stores raw provider-specific fields as unstructured data. This ensures lossless round-tripping of unknown parameters while keeping the core API provider-neutral. Additionally, the preservation field maintains exact original payloads keyed by source format for same-format replay scenarios.
What is the difference between the extensions and preservation fields in the protocol types?
The extensions field captures provider-specific parameters that lack first-class Rust struct representations but are still semantically understood as key-value pairs. The preservation field, defined as PreservationMetadata, stores the entire original request or response payload in its exact source format (identified by FormatId from crates/protocol/src/format.rs), allowing Switchyard to replay traffic without re-encoding or losing formatting nuances.
Where are the protocol types re-exported for use in other Switchyard crates?
All public protocol types are re-exported from [crates/protocol/src/lib.rs](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), making them accessible as protocol::LlmRequest and protocol::AggLlmResponse to other components within the workspace. The FormatId type used by preservation metadata resides in crates/protocol/src/format.rs, while transport wrapping is handled in crates/protocol/src/envelope.rs.
How does AggLlmResponse normalize different providers' token usage reporting?
The AggLlmResponse includes a usage field containing a Usage struct that standardizes token accounting across providers. It defines input_tokens, output_tokens, total_tokens, and detailed cache usage via InputCacheUsage. This normalization allows Switchyard to perform consistent cost tracking and rate limiting regardless of whether the backend uses OpenAI's prompt_tokens, Anthropic's input_tokens, or other vendor-specific nomenclature.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →