How Switchyard Resolves Route IDs to Targets for OpenAI-Compatible Endpoints

The Switchyard server extracts the model identifier from incoming POST requests, queries the internal Runner against the TOML configuration to retrieve a Route object, executes the configured routing algorithm to select a specific target, and injects the chosen target's model ID into the final HTTP response.

NVIDIA-NeMo/Switchyard is an intelligent request router that provides OpenAI-compatible API endpoints. When clients send requests to /v1/chat/completions, /v1/messages, or /v1/responses, the server must map the abstract route ID provided in the JSON payload to a concrete backend target. This resolution process involves a four-stage pipeline that transforms the incoming model string into a specific LLM client configuration.

The Route Resolution Pipeline

The server handles all three endpoints—/v1/chat/completions, /v1/messages, and /v1/responses—using identical resolution logic. The process converts the client-supplied model field into an executable routing decision through four distinct steps.

Step 1: Extract the Route ID from the Request Payload

When a POST request arrives, the HTTP handler parses the JSON body using switchyard_translation::decode_request. The decoder extracts the model field from the payload, which serves as the route ID.

For example, a request containing "model": "switchyard" passes this identifier to the resolution engine. This value corresponds to a named route defined in the server's routes.toml configuration file.

Step 2: Lookup the Route in ServerState

The ServerState struct maintains a Runner instance that serves as the core router. The handler invokes route_for_model to query this runner:

fn route_for_model(&self, model: &str) -> Option<&Route> {
    self.runner.route(model)
}

This method, located in crates/switchyard-server/src/lib.rs at lines 21-24, returns an optional reference to a Route struct. The Runner has pre-loaded the TOML configuration into a hash map that associates model IDs with their corresponding Route definitions, which include target lists and routing algorithm specifications.

Step 3: Execute the Routing Algorithm

Once the system retrieves the Route object, it executes the configured algorithm via Algorithm::run_stream. This produces a RoutingOutcome struct containing the selected target and any fallback configurations.

The ServerState::decision_response method (lines 26-47 in lib.rs) transforms this outcome into a DecisionResponse:

fn decision_response(&self,
                     route_model: &ModelId,
                     outcome: &RoutingOutcome,
                     response: Option<Value>) -> Option<DecisionResponse> {
    let description = self.runner.describe_decision(route_model, outcome)?;
    // ...
}

This response encapsulates the selected target's model ID, the specific LLM client to use, and metadata required for request forwarding.

Step 4: Encode the Response with the Selected Target

After the routing algorithm selects a target, the server prepares the HTTP response. The into_http_response function in crates/switchyard-server/src/response.rs (lines 19-27) receives the chosen target's model ID as the served_model parameter:

pub(crate) fn into_http_response(
    response: AlgorithmResponse,
    target_format: WireFormat,
    served_model: Option<String>,
    request_extensions: ProviderExtensions,
) -> Result<HttpResponse, BoxError> { 
    // ... 
}

This function serializes the provider-neutral response into the OpenAI wire format while injecting the actual model name that serviced the request, ensuring clients receive accurate provenance information.

Code Examples in Practice

Minimal Client Request

The following curl command demonstrates how a client specifies the route ID:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "switchyard",
        "messages": [{"role": "user", "content": "Hello"}]
      }'

Under the hood, the server performs the following actions:

  • Looks up the route named switchyard in routes.toml
  • Executes the configured algorithm (e.g., stage_router)
  • Selects between targets defined under [targets.efficient] and [targets.capable]
  • Forwards the request to the chosen LLM client
  • Returns the response with the selected model's ID in the payload

Server-Side Resolution Implementation

The handler functions for all three endpoints follow this pattern:

async fn openai_chat_completions(
    State(state): State<ServerState>,
    Json(payload): Json<Value>,
) -> Result<impl IntoResponse, ServerError> {
    // Decode the request into the internal Switchyard type
    let request = decode_request(&payload, WireFormat::OpenAiChat)?;
    
    // Resolve the route ID to a Route object
    let route = state.route_for_model(&request.model)
        .ok_or_else(|| ServerError::new("unknown route"))?;
    
    // Run the routing algorithm to obtain the chosen target
    let outcome = state.runner.run(&request, route)?;
    
    // Build the HTTP response with the target's model ID
    let http_resp = into_http_response(
        outcome.response,
        WireFormat::OpenAiChat,
        Some(outcome.selected_model_id.clone()),
        request.extensions,
    )?;
    
    Ok(http_resp)
}

TOML Configuration Mapping

The routes.toml file defines the relationship between route IDs and concrete targets:

[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"

[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"

[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

In this configuration, the switchyard route references two targets. The algorithm determines which target's id (the actual model name) services each request based on the picker strategy and confidence_threshold.

Key Source Files and Components

Understanding the resolution flow requires familiarity with these specific files in the NVIDIA-NeMo/Switchyard repository:

Summary

  • The model field in incoming JSON requests serves as the route ID that drives the entire resolution process.
  • ServerState::route_for_model queries the Runner to retrieve a Route object from the TOML configuration map.
  • The routing algorithm produces a RoutingOutcome that specifies which target handles the request, encapsulated later in a DecisionResponse.
  • into_http_response injects the selected target's model ID as served_model into the final HTTP response, ensuring accurate client-side tracking.
  • All three endpoints—/v1/chat/completions, /v1/messages, and /v1/responses—share identical route resolution logic, differing only in their request/response serialization formats.

Frequently Asked Questions

How does Switchyard handle unknown route IDs?

When route_for_model cannot find a matching entry in the Runner's configuration map, it returns None. The handler typically converts this into a ServerError with an "unknown route" message, returning a 400-level HTTP status to the client. This prevents requests from proceeding to the routing algorithm stage without a valid configuration.

Can the same route ID resolve to different targets for different requests?

Yes. The resolution process is dynamic. While the Route object remains constant for a given ID, the routing algorithm (e.g., stage_router with an efficient_first picker) evaluates each request independently. Based on confidence thresholds, latency, or other heuristics defined in the TOML configuration, the algorithm may select different targets from the same route definition for different incoming requests.

What is the difference between a route ID and a target ID?

The route ID (specified in the client request's model field) corresponds to the [routes.*] sections in TOML and represents the abstract routing strategy. The target ID (defined in [targets.*] sections) represents the concrete LLM model identifier (e.g., anthropic/claude-opus-4.8) that ultimately services the request. The resolution process maps the former to the latter through the configured algorithm.

Where does the final model name in the HTTP response originate?

The model name returned to the client originates from the target's id field in the TOML configuration. When into_http_response processes the AlgorithmResponse, it receives the selected target's identifier as the served_model parameter (lines 19-27 in response.rs). This value overrides any placeholder model names, ensuring the response accurately reflects which backend LLM processed the request.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →