Switchyard Request Lifecycle: From TCP Ingress to LLM Response
Switchyard processes incoming LLM requests through an Axum-based HTTP pipeline that timestamps ingress, resolves routes via libsy algorithms, forwards to downstream clients, and serializes responses with wire-format-specific encoding.
The NVIDIA-NeMo/Switchyard repository implements a high-performance routing layer for Large Language Model (LLM) traffic. Understanding the Switchyard request lifecycle is essential for operators tuning latency, debugging routing decisions, or extending the platform with custom algorithms. The entire flow spans from TCP socket binding to final JSON serialization, leveraging Rust’s Axum framework for async request handling.
Server Initialization and TCP Binding
The lifecycle begins when the binary entry point in crates/switchyard-server/src/main.rs invokes BoundServer::bind from crates/switchyard-server/src/lib.rs. This function binds a TCP socket and constructs the Axum router via build_switchyard_router, registering supported endpoints such as /v1/chat/completions and /v1/messages. The server configuration, parsed in crates/switchyard-server/src/config.rs, loads TOML definitions for routes, target clients, and capabilities before the server accepts traffic.
Ingress Timing and Middleware Processing
Upon accepting a connection, the stamp_request_start middleware inserts a RequestStart(Instant) into the request’s extensions. Located in crates/switchyard-server/src/lib.rs, this middleware enables downstream components to calculate total request latency by capturing the exact moment the HTTP request enters the Switchyard stack.
This timestamp persists through the entire async call chain, eventually feeding into the usage_metrics::observe call that records final duration statistics.
Endpoint Dispatch and Route Resolution
The Axum router dispatches requests to format-specific handlers—openai_chat_completions, anthropic_messages, or openai_responses—which all forward to the central handle_endpoint function. This handler, defined in crates/switchyard-server/src/lib.rs, orchestrates the critical resolve_route phase.
The resolve_route function performs four key operations:
- Decodes the JSON body into a
switchyard_protocol::Requestusingdecode_requestfromcrates/switchyard-translation/src/lib.rs. - Validates that the
modelfield is non-empty. - Looks up the model in
ServerState.routesviaroute_for_model. - Ensures the caller’s authentication kind matches the expected wire format.
Algorithm Execution and Downstream Calls
Once the route resolves, handle_llm_request takes control. This function instantiates a stats_observer and invokes switchyard_llm_client::run, which ultimately calls libsy::drive with the algorithm attached to the resolved route.
The routing algorithm, defined in crates/libsy/src/core/algorithm.rs, may issue classifier or judge calls during execution. Each dependency call routes through serve_decision_dependency, which forwards requests to the appropriate downstream LLM client via client.call in crates/libsy-llm-client/src/client.rs. This design allows recursive routing decisions where intermediate models evaluate content before selecting the final target.
Response Aggregation and Serialization
When the algorithm completes, it returns a RoutingOutcome containing the selected model ID and raw LlmResponse. The system passes these to usage_metrics::observe, which records latency, token usage, and optional routing-log entries for observability.
The into_http_response function in crates/switchyard-server/src/response.rs converts the internal LlmResponse into the wire-format-specific HTTP response. For OpenAI-compatible endpoints, this produces standard JSON payloads; for Anthropic endpoints, it adapts to the Messages API schema. The function also injects the x-model-router-selected-model header to expose routing decisions to the caller.
Error Handling and Final Delivery
Any error propagated through the call stack wraps into an ApiError conforming to either OpenAI or Anthropic error schemas. The render_error_response function in crates/switchyard-server/src/lib.rs renders these into appropriate HTTP status codes and JSON bodies.
A request-log middleware captures the final status code, duration (calculated from the initial RequestStart extension), and any error messages before Axum sends the response through the TCP socket. This completes the HTTP round-trip.
Practical Examples
To run a local Switchyard server using the Python bindings exposed in switchyard_rust/server.py:
from switchyard_rust import Server
# Load server from a TOML config that defines routes and downstream clients
srv = Server("examples/routes.toml", port=4000)
print(f"Listening on http://localhost:{srv.port}")
# The server runs until the process exits (or you call `srv.close()`)
Send an OpenAI-format chat completion request:
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "my-route",
"messages": [{"role":"user","content":"Say hello"}]
}'
Or use the Anthropic Messages API format:
curl -X POST http://localhost:4000/v1/messages \
-H "Content-Type: application/json" \
-d '{
"model": "my-anthropic-route",
"messages": [{"role":"user","content":"Explain quantum entanglement"}]
}'
Inspect routing statistics via the built-in endpoint:
curl http://localhost:4000/v1/stats | jq .
Summary
- Server startup binds TCP sockets and registers Axum routes via
BoundServer::bindincrates/switchyard-server/src/lib.rs, loading configuration fromcrates/switchyard-server/src/config.rs. - Request timing starts with the
stamp_request_startmiddleware, which inserts anInstantinto request extensions for latency tracking. - Route resolution occurs in
resolve_route, validating models againstServerState.routesand decoding JSON viadecode_requestin the translation layer. - Algorithm execution runs through
libsy::drive, potentially invoking downstream clients recursively viaserve_decision_dependencyandcrates/libsy-llm-client/src/client.rs. - Response serialization uses
into_http_responseto generate wire-format-specific JSON and adds thex-model-router-selected-modelheader. - Error handling wraps failures in
ApiErrorobjects rendered byrender_error_response, with final logging capturing status codes and durations.
Frequently Asked Questions
What Axum middleware does Switchyard use for request timing?
Switchyard uses the stamp_request_start middleware located in crates/switchyard-server/src/lib.rs. This middleware inserts a RequestStart(Instant) into the request extensions immediately upon ingress, enabling precise latency calculation from TCP acceptance to response serialization.
How does Switchyard validate incoming model names?
During the resolve_route phase in crates/switchyard-server/src/lib.rs, Switchyard validates that the model field is non-empty and then performs a lookup via route_for_model against the ServerState.routes map. If the model is undefined or the authentication kind mismatches the wire format, the request rejects before reaching the routing algorithm.
Where is the selected model ID exposed in the response?
The into_http_response function in crates/switchyard-server/src/response.rs attaches the selected model ID as the x-model-router-selected-model HTTP header. This header appears in the final HTTP response regardless of whether the wire format is OpenAI, Anthropic, or OpenAI-compatible responses.
How are routing algorithm decisions executed?
The handle_llm_request function invokes switchyard_llm_client::run, which delegates to libsy::drive with the route’s configured algorithm. The algorithm may call intermediate models (classifiers or judges) through serve_decision_dependency, which uses crates/libsy-llm-client/src/client.rs to forward requests downstream before returning a final RoutingOutcome.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →