How Switchyard Resolves Route IDs to Targets for OpenAI-Compatible Endpoints
The Switchyard server extracts the model identifier from incoming POST requests, queries the internal Runner against the TOML configuration to retrieve a Route object, executes the configured routing algorithm to select a specific target, and injects the chosen target's model ID into the final HTTP response.
NVIDIA-NeMo/Switchyard is an intelligent request router that provides OpenAI-compatible API endpoints. When clients send requests to /v1/chat/completions, /v1/messages, or /v1/responses, the server must map the abstract route ID provided in the JSON payload to a concrete backend target. This resolution process involves a four-stage pipeline that transforms the incoming model string into a specific LLM client configuration.
The Route Resolution Pipeline
The server handles all three endpoints—/v1/chat/completions, /v1/messages, and /v1/responses—using identical resolution logic. The process converts the client-supplied model field into an executable routing decision through four distinct steps.
Step 1: Extract the Route ID from the Request Payload
When a POST request arrives, the HTTP handler parses the JSON body using switchyard_translation::decode_request. The decoder extracts the model field from the payload, which serves as the route ID.
For example, a request containing "model": "switchyard" passes this identifier to the resolution engine. This value corresponds to a named route defined in the server's routes.toml configuration file.
Step 2: Lookup the Route in ServerState
The ServerState struct maintains a Runner instance that serves as the core router. The handler invokes route_for_model to query this runner:
fn route_for_model(&self, model: &str) -> Option<&Route> {
self.runner.route(model)
}
This method, located in crates/switchyard-server/src/lib.rs at lines 21-24, returns an optional reference to a Route struct. The Runner has pre-loaded the TOML configuration into a hash map that associates model IDs with their corresponding Route definitions, which include target lists and routing algorithm specifications.
Step 3: Execute the Routing Algorithm
Once the system retrieves the Route object, it executes the configured algorithm via Algorithm::run_stream. This produces a RoutingOutcome struct containing the selected target and any fallback configurations.
The ServerState::decision_response method (lines 26-47 in lib.rs) transforms this outcome into a DecisionResponse:
fn decision_response(&self,
route_model: &ModelId,
outcome: &RoutingOutcome,
response: Option<Value>) -> Option<DecisionResponse> {
let description = self.runner.describe_decision(route_model, outcome)?;
// ...
}
This response encapsulates the selected target's model ID, the specific LLM client to use, and metadata required for request forwarding.
Step 4: Encode the Response with the Selected Target
After the routing algorithm selects a target, the server prepares the HTTP response. The into_http_response function in crates/switchyard-server/src/response.rs (lines 19-27) receives the chosen target's model ID as the served_model parameter:
pub(crate) fn into_http_response(
response: AlgorithmResponse,
target_format: WireFormat,
served_model: Option<String>,
request_extensions: ProviderExtensions,
) -> Result<HttpResponse, BoxError> {
// ...
}
This function serializes the provider-neutral response into the OpenAI wire format while injecting the actual model name that serviced the request, ensuring clients receive accurate provenance information.
Code Examples in Practice
Minimal Client Request
The following curl command demonstrates how a client specifies the route ID:
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "switchyard",
"messages": [{"role": "user", "content": "Hello"}]
}'
Under the hood, the server performs the following actions:
- Looks up the route named
switchyardinroutes.toml - Executes the configured algorithm (e.g.,
stage_router) - Selects between targets defined under
[targets.efficient]and[targets.capable] - Forwards the request to the chosen LLM client
- Returns the response with the selected model's ID in the payload
Server-Side Resolution Implementation
The handler functions for all three endpoints follow this pattern:
async fn openai_chat_completions(
State(state): State<ServerState>,
Json(payload): Json<Value>,
) -> Result<impl IntoResponse, ServerError> {
// Decode the request into the internal Switchyard type
let request = decode_request(&payload, WireFormat::OpenAiChat)?;
// Resolve the route ID to a Route object
let route = state.route_for_model(&request.model)
.ok_or_else(|| ServerError::new("unknown route"))?;
// Run the routing algorithm to obtain the chosen target
let outcome = state.runner.run(&request, route)?;
// Build the HTTP response with the target's model ID
let http_resp = into_http_response(
outcome.response,
WireFormat::OpenAiChat,
Some(outcome.selected_model_id.clone()),
request.extensions,
)?;
Ok(http_resp)
}
TOML Configuration Mapping
The routes.toml file defines the relationship between route IDs and concrete targets:
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"
[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"
[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
In this configuration, the switchyard route references two targets. The algorithm determines which target's id (the actual model name) services each request based on the picker strategy and confidence_threshold.
Key Source Files and Components
Understanding the resolution flow requires familiarity with these specific files in the NVIDIA-NeMo/Switchyard repository:
crates/switchyard-server/src/lib.rs: ContainsServerState, theroute_for_modellookup implementation (lines 21-24), and thedecision_responsebuilder (lines 26-47).crates/switchyard-server/src/response.rs: Implementsinto_http_response(lines 19-27), which encodes the algorithm's output into the OpenAI/Anthropic wire format while inserting the selected model name.routes.toml: The user-provided configuration file that maps route IDs to target definitions and routing algorithms.crates/switchyard-runner/src/route.rs: Implements theRunner::routemethod that loads the TOML schema and returnsRoutestructs to the server.crates/switchyard-server/src/cli.rs: Parses command-line arguments and initializes the server with the TOML configuration.crates/switchyard-server/src/metrics.rs: Records per-request metrics, including which target served each call for observability purposes.
Summary
- The model field in incoming JSON requests serves as the route ID that drives the entire resolution process.
ServerState::route_for_modelqueries theRunnerto retrieve aRouteobject from the TOML configuration map.- The routing algorithm produces a
RoutingOutcomethat specifies which target handles the request, encapsulated later in aDecisionResponse. into_http_responseinjects the selected target's model ID asserved_modelinto the final HTTP response, ensuring accurate client-side tracking.- All three endpoints—
/v1/chat/completions,/v1/messages, and/v1/responses—share identical route resolution logic, differing only in their request/response serialization formats.
Frequently Asked Questions
How does Switchyard handle unknown route IDs?
When route_for_model cannot find a matching entry in the Runner's configuration map, it returns None. The handler typically converts this into a ServerError with an "unknown route" message, returning a 400-level HTTP status to the client. This prevents requests from proceeding to the routing algorithm stage without a valid configuration.
Can the same route ID resolve to different targets for different requests?
Yes. The resolution process is dynamic. While the Route object remains constant for a given ID, the routing algorithm (e.g., stage_router with an efficient_first picker) evaluates each request independently. Based on confidence thresholds, latency, or other heuristics defined in the TOML configuration, the algorithm may select different targets from the same route definition for different incoming requests.
What is the difference between a route ID and a target ID?
The route ID (specified in the client request's model field) corresponds to the [routes.*] sections in TOML and represents the abstract routing strategy. The target ID (defined in [targets.*] sections) represents the concrete LLM model identifier (e.g., anthropic/claude-opus-4.8) that ultimately services the request. The resolution process maps the former to the latter through the configured algorithm.
Where does the final model name in the HTTP response originate?
The model name returned to the client originates from the target's id field in the TOML configuration. When into_http_response processes the AlgorithmResponse, it receives the selected target's identifier as the served_model parameter (lines 19-27 in response.rs). This value overrides any placeholder model names, ensuring the response accurately reflects which backend LLM processed the request.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →