# How Switchyard Resolves Route IDs to Targets for OpenAI-Compatible Endpoints

> Learn how Switchyard resolves route IDs to targets for OpenAI-compatible endpoints like v1 chat completions. Discover the routing algorithm and target selection process.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-12

---

**The Switchyard server extracts the model identifier from incoming POST requests, queries the internal Runner against the TOML configuration to retrieve a Route object, executes the configured routing algorithm to select a specific target, and injects the chosen target's model ID into the final HTTP response.**

NVIDIA-NeMo/Switchyard is an intelligent request router that provides OpenAI-compatible API endpoints. When clients send requests to `/v1/chat/completions`, `/v1/messages`, or `/v1/responses`, the server must map the abstract route ID provided in the JSON payload to a concrete backend target. This resolution process involves a four-stage pipeline that transforms the incoming model string into a specific LLM client configuration.

## The Route Resolution Pipeline

The server handles all three endpoints—`/v1/chat/completions`, `/v1/messages`, and `/v1/responses`—using identical resolution logic. The process converts the client-supplied model field into an executable routing decision through four distinct steps.

### Step 1: Extract the Route ID from the Request Payload

When a POST request arrives, the HTTP handler parses the JSON body using `switchyard_translation::decode_request`. The decoder extracts the `model` field from the payload, which serves as the route ID.

For example, a request containing `"model": "switchyard"` passes this identifier to the resolution engine. This value corresponds to a named route defined in the server's [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) configuration file.

### Step 2: Lookup the Route in ServerState

The `ServerState` struct maintains a `Runner` instance that serves as the core router. The handler invokes `route_for_model` to query this runner:

```rust
fn route_for_model(&self, model: &str) -> Option<&Route> {
    self.runner.route(model)
}

```

This method, located in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) at lines 21-24, returns an optional reference to a `Route` struct. The `Runner` has pre-loaded the TOML configuration into a hash map that associates model IDs with their corresponding `Route` definitions, which include target lists and routing algorithm specifications.

### Step 3: Execute the Routing Algorithm

Once the system retrieves the `Route` object, it executes the configured algorithm via `Algorithm::run_stream`. This produces a `RoutingOutcome` struct containing the selected target and any fallback configurations.

The `ServerState::decision_response` method (lines 26-47 in [`lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/lib.rs)) transforms this outcome into a `DecisionResponse`:

```rust
fn decision_response(&self,
                     route_model: &ModelId,
                     outcome: &RoutingOutcome,
                     response: Option<Value>) -> Option<DecisionResponse> {
    let description = self.runner.describe_decision(route_model, outcome)?;
    // ...
}

```

This response encapsulates the **selected target's model ID**, the specific LLM client to use, and metadata required for request forwarding.

### Step 4: Encode the Response with the Selected Target

After the routing algorithm selects a target, the server prepares the HTTP response. The `into_http_response` function in [`crates/switchyard-server/src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/response.rs) (lines 19-27) receives the chosen target's model ID as the `served_model` parameter:

```rust
pub(crate) fn into_http_response(
    response: AlgorithmResponse,
    target_format: WireFormat,
    served_model: Option<String>,
    request_extensions: ProviderExtensions,
) -> Result<HttpResponse, BoxError> { 
    // ... 
}

```

This function serializes the provider-neutral response into the OpenAI wire format while injecting the actual model name that serviced the request, ensuring clients receive accurate provenance information.

## Code Examples in Practice

### Minimal Client Request

The following curl command demonstrates how a client specifies the route ID:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "switchyard",
        "messages": [{"role": "user", "content": "Hello"}]
      }'

```

Under the hood, the server performs the following actions:

- Looks up the route named `switchyard` in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml)
- Executes the configured algorithm (e.g., `stage_router`)
- Selects between targets defined under `[targets.efficient]` and `[targets.capable]`
- Forwards the request to the chosen LLM client
- Returns the response with the selected model's ID in the payload

### Server-Side Resolution Implementation

The handler functions for all three endpoints follow this pattern:

```rust
async fn openai_chat_completions(
    State(state): State<ServerState>,
    Json(payload): Json<Value>,
) -> Result<impl IntoResponse, ServerError> {
    // Decode the request into the internal Switchyard type
    let request = decode_request(&payload, WireFormat::OpenAiChat)?;
    
    // Resolve the route ID to a Route object
    let route = state.route_for_model(&request.model)
        .ok_or_else(|| ServerError::new("unknown route"))?;
    
    // Run the routing algorithm to obtain the chosen target
    let outcome = state.runner.run(&request, route)?;
    
    // Build the HTTP response with the target's model ID
    let http_resp = into_http_response(
        outcome.response,
        WireFormat::OpenAiChat,
        Some(outcome.selected_model_id.clone()),
        request.extensions,
    )?;
    
    Ok(http_resp)
}

```

### TOML Configuration Mapping

The [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file defines the relationship between route IDs and concrete targets:

```toml
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

```

In this configuration, the `switchyard` route references two targets. The algorithm determines which target's `id` (the actual model name) services each request based on the `picker` strategy and `confidence_threshold`.

## Key Source Files and Components

Understanding the resolution flow requires familiarity with these specific files in the NVIDIA-NeMo/Switchyard repository:

- **[`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs)**: Contains `ServerState`, the `route_for_model` lookup implementation (lines 21-24), and the `decision_response` builder (lines 26-47).
- **[`crates/switchyard-server/src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/response.rs)**: Implements `into_http_response` (lines 19-27), which encodes the algorithm's output into the OpenAI/Anthropic wire format while inserting the selected model name.
- **[`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml)**: The user-provided configuration file that maps route IDs to target definitions and routing algorithms.
- **[`crates/switchyard-runner/src/route.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/route.rs)**: Implements the `Runner::route` method that loads the TOML schema and returns `Route` structs to the server.
- **[`crates/switchyard-server/src/cli.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/cli.rs)**: Parses command-line arguments and initializes the server with the TOML configuration.
- **[`crates/switchyard-server/src/metrics.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/metrics.rs)**: Records per-request metrics, including which target served each call for observability purposes.

## Summary

- The **model field** in incoming JSON requests serves as the route ID that drives the entire resolution process.
- `ServerState::route_for_model` queries the `Runner` to retrieve a `Route` object from the TOML configuration map.
- The routing algorithm produces a `RoutingOutcome` that specifies which target handles the request, encapsulated later in a `DecisionResponse`.
- `into_http_response` injects the selected target's model ID as `served_model` into the final HTTP response, ensuring accurate client-side tracking.
- All three endpoints—`/v1/chat/completions`, `/v1/messages`, and `/v1/responses`—share identical route resolution logic, differing only in their request/response serialization formats.

## Frequently Asked Questions

### How does Switchyard handle unknown route IDs?

When `route_for_model` cannot find a matching entry in the Runner's configuration map, it returns `None`. The handler typically converts this into a `ServerError` with an "unknown route" message, returning a 400-level HTTP status to the client. This prevents requests from proceeding to the routing algorithm stage without a valid configuration.

### Can the same route ID resolve to different targets for different requests?

Yes. The resolution process is dynamic. While the `Route` object remains constant for a given ID, the routing algorithm (e.g., `stage_router` with an `efficient_first` picker) evaluates each request independently. Based on confidence thresholds, latency, or other heuristics defined in the TOML configuration, the algorithm may select different targets from the same route definition for different incoming requests.

### What is the difference between a route ID and a target ID?

The **route ID** (specified in the client request's `model` field) corresponds to the `[routes.*]` sections in TOML and represents the abstract routing strategy. The **target ID** (defined in `[targets.*]` sections) represents the concrete LLM model identifier (e.g., `anthropic/claude-opus-4.8`) that ultimately services the request. The resolution process maps the former to the latter through the configured algorithm.

### Where does the final model name in the HTTP response originate?

The model name returned to the client originates from the target's `id` field in the TOML configuration. When `into_http_response` processes the `AlgorithmResponse`, it receives the selected target's identifier as the `served_model` parameter (lines 19-27 in [`response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/response.rs)). This value overrides any placeholder model names, ensuring the response accurately reflects which backend LLM processed the request.