# How the Switchyard Server Resolves LLM Routes: A Deep Dive into the Routing Pipeline

> Discover how the Switchyard server resolves LLM routes via a three-stage pipeline. Learn about identifier extraction, route lookup, and upstream target resolution from NVIDIA-NeMo/Switchyard.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-21

---

**The Switchyard server resolves LLM routes through a three-stage pipeline that extracts the model identifier from request metadata, looks up the corresponding route entry in a B-tree map, and resolves concrete upstream targets using the server's routing configuration.**

The NVIDIA-NeMo/Switchyard repository provides a high-performance routing layer for LLM requests. Understanding how the Switchyard server resolves LLM routes requires examining the internal flow from HTTP request to concrete upstream target selection.

## The Three-Stage Route Resolution Pipeline

The resolution process follows distinct stages defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).

### Stage 1: Extracting the Model Identifier from Request Metadata

Incoming requests for completion or chat endpoints must include a `"model"` field in the JSON body. This field names the **synthetic route model** configured in the server.

The server also parses HTTP headers via `metadata_from_headers` to capture additional metadata. For example, the header `x-model-router-selected-model` can influence routing decisions. This metadata extraction occurs in the request handling layer before route resolution begins.

### Stage 2: Looking Up the Route Entry in ServerState

The `ServerState` struct maintains a `BTreeMap<ModelId, RouteEntry>` (accessible via `self.routes`). The `resolve_route` function calls `state.route_for_model(&model)` to retrieve the corresponding `RouteEntry`.

If the model identifier is unknown, the resolver returns a **404-style error response** immediately. This lookup operation is defined around lines 86-88 in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs).

### Stage 3: Resolving Concrete Upstream Targets

Once a `RouteEntry` is found, the server transforms the JSON body into a strongly-typed `switchyard_protocol::Request` using `decode_request` from the translation layer. The routing algorithm selects a model (and potential fallbacks), which are then mapped to concrete upstream targets using `config.routing_target_names`.

Each target is resolved through the per-target `ClientRouter` stored in the `RouteEntry`. The final resolved information is packaged in a `DecisionResponse`, which the server returns when callers query the debug endpoint.

## Core Resolver Implementation

The **core resolver** lives in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) around line 858. Its signature is:

```rust
fn resolve_route(
    state: &ServerState,
    metadata: Metadata,
    body: Value,
    wire_format: WireFormat,
) -> Result<(Arc<RouteEntry>, Request), Response>

```

Inside this function, the server performs these steps:

1. Extracts the `"model"` field from `body`, returning an *invalid request* error if missing.
2. Calls `state.route_for_model(&model)` to fetch the `RouteEntry`.
3. Maps the algorithm's selected `ModelId` to concrete upstream targets using the server's configuration.
4. Returns a tuple containing the route entry and the fully-typed request for the algorithm to drive.

If any step fails—whether from an unknown model, missing target configuration, or malformed JSON—the resolver returns an appropriate HTTP response with a clear error payload (e.g., `invalid_body_error`, `server_error`, or an *unknown route* response).

## Key Source Files and Components

Understanding the route resolution requires familiarity with these specific files:

- [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs): Contains `ServerState`, the `routes` map, and the `resolve_route` logic.
- [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs): Provides `decode_request` and `encode_aggregated_response` for converting between JSON payloads and internal `Request` types.
- [`crates/protocol/src/model_id.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/model_id.rs): Defines `ModelId`, the strongly-typed identifier used as keys in the routing map.
- [`crates/protocol/src/metadata.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/metadata.rs): Parses request-level metadata from headers that can affect routing decisions.
- [`crates/switchyard-py/src/server_bindings.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/server_bindings.rs): Exposes the Rust server to Python via PyO3, allowing Python callers to leverage the same route-resolution logic.

## Practical Examples

### Python Client Requesting a Synthetic Model

To trigger route resolution, clients must specify the synthetic route name in the request body:

```python
import httpx

url = "http://localhost:4000/v1/chat/completions"
payload = {
    "model": "my-synthetic-route",  # Synthetic route name from server config

    "messages": [{"role": "user", "content": "Hello"}],
    "temperature": 0.7,
}
resp = httpx.post(url, json=payload)
print(resp.json())

```

### Inspecting Resolved Routing Decisions

Switchyard provides a debug endpoint to inspect routing decisions without executing the full completion:

```bash
curl -s http://localhost:4000/v1/models?model=my-synthetic-route \
     -H "Accept: application/json"

```

The response contains the `selected` downstream target, the wire format (e.g., `openai_chat`, `anthropic_messages`), and the base URL for the actual LLM call.

## Summary

- The Switchyard server uses a three-stage pipeline to resolve LLM routes: metadata extraction, route lookup, and target resolution.
- `ServerState` maintains a `BTreeMap` mapping `ModelId` to `RouteEntry` for O(log n) lookups.
- The `resolve_route` function in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) orchestrates the resolution and provides detailed error handling.
- Requests must include a `"model"` field naming the synthetic route configured in the server.
- Routing decisions can be inspected via the `/v1/models` debug endpoint without sending completion requests.

## Frequently Asked Questions

### What happens if the model name is not found in the routes map?

If `state.route_for_model(&model)` fails to find the identifier in the `BTreeMap<ModelId, RouteEntry>`, the resolver returns a **404-style error response** with a clear error payload indicating an unknown route.

### How does Switchyard handle request metadata from headers?

The server parses HTTP headers using `metadata_from_headers` (defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) around lines 85-90). This captures metadata such as `x-model-router-selected-model`, which can influence routing decisions alongside the body content.

### What is the role of RouteEntry in the resolution process?

`RouteEntry` stores the per-route configuration including the `ClientRouter` used to resolve concrete upstream targets. After the server looks up the route by `ModelId`, the `RouteEntry` provides the `target_clients` mapping and configuration needed to translate the request into an actual upstream LLM call.

### Can I inspect routing decisions without sending a completion request?

Yes. Switchyard exposes a debug endpoint at `/v1/models` that returns a `DecisionResponse` showing the selected model, fallback models, wire format, and target URLs. Query this endpoint with the `model` parameter to see how the Switchyard server resolves LLM routes for specific synthetic model names.