# Switchyard RoutingOutcome Structure: How Model Selection Drives the Answer Call

> Understand Switchyard's RoutingOutcome structure. Learn how selected model IDs, request, response, and metadata influence the answer call for efficient routing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**The `RoutingOutcome` structure in NVIDIA-NeMo/Switchyard is the core data object returned by routing algorithms that encapsulates selected model IDs, the original request, an optional pre-fetched response, and metadata, determining whether the `answer` call returns immediately or triggers a model invocation.**

Switchyard is an open-source LLM routing framework developed by NVIDIA. The `RoutingOutcome` structure defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) serves as the contract between routing algorithms and the execution engine, carrying all necessary context for the `answer` call to fulfill client requests.

## RoutingOutcome Structure Definition

In [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), the structure is defined as a plain data container with public fields:

```rust
pub struct RoutingOutcome {
    /// The ID(s) of the model(s) that were finally selected.
    pub selected_model_ids: Vec<String>,
    /// The original request that entered the routing pipeline.
    pub request: Request,
    /// The response generated by the selected model (if the algorithm
    /// performed a call).  This field is `None` for a "route-to" outcome
    /// that has not yet called a model.
    pub response: Option<Response>,
    /// Arbitrary metadata that algorithms can attach for observability
    /// or downstream processing (e.g., latency, cost, token counts).
    pub metadata: HashMap<String, String>,
}

```

The structure provides two helper constructors used by algorithms: `route_to()` for outcomes requiring deferred model invocation, and `answered()` for outcomes that already contain a cached or pre-fetched response.

## How RoutingOutcome Fields Control the Answer Workflow

The `answer` call consumes a `RoutingOutcome` to determine whether to return a response immediately or invoke the selected model. Each field carries specific implications for this decision.

### selected_model_ids and Target Selection

The `selected_model_ids` field contains a `Vec<String>` listing the model identifiers chosen by the routing algorithm. According to the implementation in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs), the `answer` call uses this list to know exactly which model (or fallback chain) to query.

If the algorithm selected multiple models for fallback, the driver attempts them in order until one succeeds. This field enables downstream components to record which model was used for billing, auditing, or performance tracking.

### request Propagation

The `request` field stores the original `Request` object that entered the router, including the prompt and generation parameters. The `answer` call passes this unchanged to the model driver, ensuring the model receives the exact payload the client sent. This immutability guarantee prevents routing logic from unintentionally modifying the payload during the selection phase.

### response and Lazy Evaluation

The `response` field is an `Option<Response>` that determines the execution path:

- **When `Some(Response)`**: The algorithm performed a pre-fetch (e.g., cache lookup). The `answer` call returns this response immediately without invoking a model.
- **When `None`**: The outcome represents a "route-to" directive. The `answer` call triggers the driver to invoke the model(s) listed in `selected_model_ids` using the stored `request`, then populates this field with the result.

### metadata for Observability

The `metadata` field is a `HashMap<String, String>` for arbitrary key-value pairs. Algorithms attach diagnostic data such as routing latency, chosen fallback paths, cost estimates, or token counts. The `answer` call propagates this metadata to HTTP headers, logging systems, or Prometheus metrics, giving operators insight into why specific routing decisions were made.

## RoutingOutcome in Action: Code Examples

The following examples demonstrate how algorithms construct outcomes and how the system consumes them.

### Rust Implementation Pattern

Algorithms in [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs) construct outcomes using helper methods:

```rust
use switchyard_libsy::{RoutingOutcome, Request, Response};

// Example: Algorithm returning a "route-to" outcome
let outcome = RoutingOutcome::route_to(
    vec!["gpt-4o-mini".to_string()],
    original_request,
);

// The answer call handles both cached and deferred responses
async fn produce_answer(outcome: RoutingOutcome) -> Result<Response, LibsyError> {
    match outcome.response {
        // Cached response – return immediately
        Some(resp) => Ok(resp),
        
        // No response yet – invoke selected models
        None => {
            let resp = driver.call_model(
                outcome.selected_model_ids,
                outcome.request,
            ).await?;
            
            // Metadata can be updated post-invocation
            let mut final_outcome = outcome;
            final_outcome.response = Some(resp.clone());
            Ok(resp)
        }
    }
}

```

### Python Bindings Usage

The Python interface exposed in [`crates/switchyard-py/src/libsy_bindings.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/libsy_bindings.rs) mirrors the Rust API:

```python
from switchyard.libsy import RoutingOutcome, answer

# Create a routing outcome targeting a specific model

outcome = RoutingOutcome.route_to(
    selected_model_ids=["gpt-4o-mini"],
    request=my_request,
)

# The answer function determines whether to call the model

# or return a cached response

final_response = answer(outcome)

print(final_response.text)
print(outcome.metadata.get("routing_latency"))

```

## Key Source Files and Implementation Details

| File | Purpose |
|------|---------|
| [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) | Defines the `RoutingOutcome` struct and its constructors (`route_to`, `answered`). |
| [`crates/libsy/src/algorithms/fall_through.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/fall_through.rs) | Demonstrates algorithm implementation building outcomes with fallback chains. |
| [`crates/switchyard-py/src/libsy_bindings.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/libsy_bindings.rs) | Python wrapper exposing `RoutingOutcome` to the `switchyard` PyPI package. |
| [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs) | Contains the driver logic that consumes `RoutingOutcome` and executes model calls when `response` is `None`. |

## Summary

- The `RoutingOutcome` struct in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) is the standardized return type for all Switchyard routing algorithms.
- **`selected_model_ids`** determines which model(s) the driver invokes, supporting single-model selection and fallback chains.
- **`request`** carries the original client payload unchanged through the routing pipeline to ensure model inputs remain consistent.
- **`response`** controls execution flow: `Some` enables immediate returns, while `None` triggers lazy model invocation during the `answer` call.
- **`metadata`** provides a flexible mechanism for algorithms to attach observability data that downstream systems consume for monitoring and billing.

## Frequently Asked Questions

### What is the RoutingOutcome structure in Switchyard?

The `RoutingOutcome` structure is the core data object defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs) that encapsulates the result of a routing decision. It contains the selected model IDs, the original request, an optional pre-computed response, and a metadata map, serving as the contract between routing algorithms and the execution engine.

### How does the answer call use selected_model_ids?

The `answer` call passes the `selected_model_ids` vector to the model driver in [`crates/libsy-llm-client/src/run.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/run.rs). The driver attempts to invoke models in the order specified, moving to the next ID if the current model fails, enabling automatic fallback behavior without additional client logic.

### When is the response field None versus Some?

The `response` field is `Some(Response)` when the routing algorithm performed a pre-fetch operation, such as a cache hit or synthetic response generation. It is `None` when the algorithm only selected a target model (a "route-to" outcome), signaling the `answer` call to perform the actual model invocation.

### What metadata can algorithms attach to RoutingOutcome?

Algorithms can attach any string key-value pairs to the `metadata` HashMap, including routing latency measurements, cost estimates, token usage counts, or debugging information. This metadata propagates through the `answer` call to logging systems and monitoring dashboards for operational visibility.