# How the Switchyard Protocol Crate Defines Provider-Neutral Request and Response Types

> Discover how the Switchyard protocol crate defines provider-neutral request and response types with pure-Rust structs for unified LLM interactions and lossless data handling.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-08-22

---

**The `protocol` crate defines pure-Rust structs `LlmRequest` and `AggLlmResponse` that normalize LLM interactions from any provider into a provider-neutral format, enabling unified routing and lossless round-tripping through fields like `extensions` and `preservation`.**

The NVIDIA-NeMo/Switchyard inference router relies on the `protocol` crate to establish a canonical representation of LLM traffic. By mapping every OpenAI, Anthropic, or NVIDIA native payload into these normalized types, Switchyard performs cross-provider routing, A/B testing, and logging without maintaining bespoke logic for each backend.

## Core Request Type: LlmRequest

The primary request structure is defined in [[`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs) at line 103. The `LlmRequest` struct captures every concept common to modern LLM APIs while remaining agnostic to any specific provider's JSON schema.

### Key Fields and Provider-Neutral Semantics

The struct contains eleven top-level fields that cover the full spectrum of inference parameters:

- **model**: An optional model identifier that routing layers may rewrite before forwarding to a specific backend.
- **instructions**: System or developer blocks separated from conversational turns, supporting providers that distinguish system-level prompts.
- **messages**: An ordered vector of `Message` structs representing the conversation history, each containing a `Role` and vector of `ContentBlock`.
- **tools**: A list of tool definitions following a generic schema that mirrors OpenAI's function format but works across providers.
- **tool_choice**: A normalized policy enum (auto, required, none, or specific tool) that unifies various provider-specific tool selection fields.
- **sampling**: Contains `temperature`, `top_p`, and `top_k` parameters as a unified `SamplingParams` struct.
- **output**: Specifies token budgets through `max_output_tokens` and optional `response_format` constraints.
- **reasoning**: Captures provider-specific reasoning controls like `effort` levels or raw reasoning flags in a `ReasoningParams` struct.
- **stream**: A boolean flag that unifies streaming requests across all backends.
- **extensions**: A `ProviderExtensions` field holding raw provider-specific fields that lack first-class representation, ensuring lossless round-tripping.
- **preservation**: `PreservationMetadata` that stores exact request payloads keyed by source format, enabling same-format replay without re-encoding.

### Supporting Types for Content and Roles

The request relies on several supporting enums defined in the same file:

- **`Role`**: Normalizes actor identities into `System`, `Developer`, `User`, `Assistant`, and `Tool` variants.
- **`ContentBlock`**: Represents message content through variants including `Text`, `Reasoning`, `Image`, `Audio`, `Video`, `File`, `ToolCall`, `ToolResult`, `Refusal`, and a generic `Unknown` catch-all.
- **`Message`**: Combines a `Role` with a vector of `ContentBlock` elements to represent a single conversation turn.

Additional supporting structs include `ToolDefinition`, `ToolChoice`, `SamplingParams`, `OutputParams`, `ReasoningParams`, `ProviderExtensions`, and `PreservationMetadata`.

## Core Response Type: AggLlmResponse

The response counterpart is defined at line 38 of [[`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs). The `AggLlmResponse` struct provides a normalized view of LLM outputs that preserves provider-specific metadata while presenting a unified interface.

### Response Structure and Usage Tracking

The response type contains six primary fields:

- **id**: An optional identifier returned by the provider, kept verbatim for traceability.
- **model**: The model name reported by the backend, which may differ from the request's model field after routing decisions.
- **outputs**: An ordered list of `ResponseOutput` structs, each containing a role, vector of `ContentBlock`, and optional `StopReason`.
- **usage**: A normalized `Usage` struct containing `input_tokens`, `output_tokens`, `total_tokens`, cache details via `InputCacheUsage`, and reasoning token counts.
- **extensions**: Provider-specific response fields without first-class equivalents, stored in `ProviderExtensions` for lossless round-tripping.
- **preservation**: Exact response payloads keyed by source format, mirroring the request preservation mechanism.

### Stop Reasons and Output Handling

Each `ResponseOutput` within the `outputs` vector includes a `stop_reason` field using the `StopReason` enum. This enum normalizes termination conditions across providers:

- `EndTurn`: Natural completion of the response.
- `MaxTokens`: Generation halted due to token budget exhaustion.
- `ToolUse`: The model elected to invoke a tool.
- `ContentFilter`: Output blocked by safety filters.
- `Error`: Provider-reported error conditions.
- `Unknown`: Unrecognized stop conditions from novel providers.

## Architecture Benefits of Provider-Neutral Types

Because every Switchyard backend—whether a Rust server, Python façade, or external provider client—first translates native payloads into these canonical types, the system gains several architectural advantages:

- **Unified Routing**: The router evaluates `model`, `messages`, and `sampling` parameters without parsing provider-specific JSON variants.
- **Lossless Transformation**: The `extensions` and `preservation` fields guarantee that unknown provider fields survive round-trips, enabling "same-format" replay scenarios where the exact original payload must be reconstructed.
- **Cross-Provider Logging**: A single logging implementation can record `AggLlmResponse` structs with normalized `Usage` metrics regardless of whether the backend was OpenAI, Anthropic, or a local vLLM instance.

## Working with Protocol Types in Code

All structs are re-exported from the crate root in [[`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), making them accessible as `protocol::LlmRequest` and `protocol::AggLlmResponse`.

### Rust Examples

Constructing a request manually:

```rust
use switchyard_protocol::{LlmRequest, Message, Role, ContentBlock};

let req = LlmRequest {
    model: Some("mixtral-8x7b".into()),
    messages: vec![
        Message::text(Role::User, "Write a Rust hello world program."),
    ],
    ..Default::default()
};

// Serialize to JSON for transport
let json = serde_json::to_string(&req).unwrap();

```

Handling a normalized response:

```rust
use switchyard_protocol::{AggLlmResponse, StopReason};

fn handle_response(resp: AggLlmResponse) {
    if let Some(first) = resp.first_output() {
        if let Some(StopReason::ToolUse) = first.stop_reason {
            println!("Model requested a tool call");
        }
    }
}

```

### Python Bindings

When using Switchyard's Python interface, the same types are available through generated bindings:

```python
from switchyard.protocol import LlmRequest, Role, Message

# Build a simple single-turn request

request = LlmRequest.text_request(
    model="gpt-4o-mini",
    prompt="Explain the waterfall model."
)

print(request.model)                # → "gpt-4o-mini"

print(request.messages[0].role)     # → Role.User

```

## Summary

- The `protocol` crate defines `LlmRequest` (line 103) and `AggLlmResponse` (line 38) in [`crates/protocol/src/llm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/llm.rs) as the canonical provider-neutral types for all Switchyard traffic.
- `LlmRequest` unifies model selection, conversation history, tool definitions, sampling parameters, and reasoning controls while preserving unknown fields via `extensions` and `preservation`.
- `AggLlmResponse` normalizes response identifiers, outputs, token usage, and stop reasons across all LLM providers.
- Supporting types like `ContentBlock`, `Role`, and `StopReason` abstract provider-specific representations into generic enums.
- These types enable unified routing, logging, and lossless round-tripping regardless of the underlying LLM backend.

## Frequently Asked Questions

### How does the protocol crate handle provider-specific fields that don't map to standard LLM parameters?

The `LlmRequest` and `AggLlmResponse` structs include an `extensions` field of type `ProviderExtensions` that stores raw provider-specific fields as unstructured data. This ensures lossless round-tripping of unknown parameters while keeping the core API provider-neutral. Additionally, the `preservation` field maintains exact original payloads keyed by source format for same-format replay scenarios.

### What is the difference between the `extensions` and `preservation` fields in the protocol types?

The `extensions` field captures provider-specific parameters that lack first-class Rust struct representations but are still semantically understood as key-value pairs. The `preservation` field, defined as `PreservationMetadata`, stores the entire original request or response payload in its exact source format (identified by `FormatId` from [`crates/protocol/src/format.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/format.rs)), allowing Switchyard to replay traffic without re-encoding or losing formatting nuances.

### Where are the protocol types re-exported for use in other Switchyard crates?

All public protocol types are re-exported from [[`crates/protocol/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs)](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/lib.rs), making them accessible as `protocol::LlmRequest` and `protocol::AggLlmResponse` to other components within the workspace. The `FormatId` type used by preservation metadata resides in [`crates/protocol/src/format.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/format.rs), while transport wrapping is handled in [`crates/protocol/src/envelope.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/envelope.rs).

### How does AggLlmResponse normalize different providers' token usage reporting?

The `AggLlmResponse` includes a `usage` field containing a `Usage` struct that standardizes token accounting across providers. It defines `input_tokens`, `output_tokens`, `total_tokens`, and detailed cache usage via `InputCacheUsage`. This normalization allows Switchyard to perform consistent cost tracking and rate limiting regardless of whether the backend uses OpenAI's `prompt_tokens`, Anthropic's `input_tokens`, or other vendor-specific nomenclature.