# Switchyard Request Lifecycle: From TCP Ingress to LLM Response

> Explore the Switchyard request lifecycle. See how it handles LLM requests from TCP ingress to Axum pipeline routing, client forwarding, and response serialization.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-08-22

---

**Switchyard processes incoming LLM requests through an Axum-based HTTP pipeline that timestamps ingress, resolves routes via `libsy` algorithms, forwards to downstream clients, and serializes responses with wire-format-specific encoding.**

The NVIDIA-NeMo/Switchyard repository implements a high-performance routing layer for Large Language Model (LLM) traffic. Understanding the Switchyard request lifecycle is essential for operators tuning latency, debugging routing decisions, or extending the platform with custom algorithms. The entire flow spans from TCP socket binding to final JSON serialization, leveraging Rust’s Axum framework for async request handling.

## Server Initialization and TCP Binding

The lifecycle begins when the binary entry point in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs) invokes `BoundServer::bind` from [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs). This function binds a TCP socket and constructs the Axum router via `build_switchyard_router`, registering supported endpoints such as `/v1/chat/completions` and `/v1/messages`. The server configuration, parsed in [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs), loads TOML definitions for routes, target clients, and capabilities before the server accepts traffic.

## Ingress Timing and Middleware Processing

Upon accepting a connection, the `stamp_request_start` middleware inserts a `RequestStart(Instant)` into the request’s extensions. Located in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), this middleware enables downstream components to calculate total request latency by capturing the exact moment the HTTP request enters the Switchyard stack.

This timestamp persists through the entire async call chain, eventually feeding into the `usage_metrics::observe` call that records final duration statistics.

## Endpoint Dispatch and Route Resolution

The Axum router dispatches requests to format-specific handlers—`openai_chat_completions`, `anthropic_messages`, or `openai_responses`—which all forward to the central `handle_endpoint` function. This handler, defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), orchestrates the critical `resolve_route` phase.

The `resolve_route` function performs four key operations:

- Decodes the JSON body into a `switchyard_protocol::Request` using `decode_request` from [`crates/switchyard-translation/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-translation/src/lib.rs).
- Validates that the `model` field is non-empty.
- Looks up the model in `ServerState.routes` via `route_for_model`.
- Ensures the caller’s authentication kind matches the expected wire format.

## Algorithm Execution and Downstream Calls

Once the route resolves, `handle_llm_request` takes control. This function instantiates a `stats_observer` and invokes `switchyard_llm_client::run`, which ultimately calls `libsy::drive` with the algorithm attached to the resolved route.

The routing algorithm, defined in [`crates/libsy/src/core/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/core/algorithm.rs), may issue classifier or judge calls during execution. Each dependency call routes through `serve_decision_dependency`, which forwards requests to the appropriate downstream LLM client via `client.call` in [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs). This design allows recursive routing decisions where intermediate models evaluate content before selecting the final target.

## Response Aggregation and Serialization

When the algorithm completes, it returns a `RoutingOutcome` containing the selected model ID and raw `LlmResponse`. The system passes these to `usage_metrics::observe`, which records latency, token usage, and optional routing-log entries for observability.

The `into_http_response` function in [`crates/switchyard-server/src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/response.rs) converts the internal `LlmResponse` into the wire-format-specific HTTP response. For OpenAI-compatible endpoints, this produces standard JSON payloads; for Anthropic endpoints, it adapts to the Messages API schema. The function also injects the `x-model-router-selected-model` header to expose routing decisions to the caller.

## Error Handling and Final Delivery

Any error propagated through the call stack wraps into an `ApiError` conforming to either OpenAI or Anthropic error schemas. The `render_error_response` function in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs) renders these into appropriate HTTP status codes and JSON bodies.

A request-log middleware captures the final status code, duration (calculated from the initial `RequestStart` extension), and any error messages before Axum sends the response through the TCP socket. This completes the HTTP round-trip.

## Practical Examples

To run a local Switchyard server using the Python bindings exposed in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py):

```python
from switchyard_rust import Server

# Load server from a TOML config that defines routes and downstream clients

srv = Server("examples/routes.toml", port=4000)
print(f"Listening on http://localhost:{srv.port}")

# The server runs until the process exits (or you call `srv.close()`)

```

Send an OpenAI-format chat completion request:

```bash
curl -X POST http://localhost:4000/v1/chat/completions \
     -H "Content-Type: application/json" \
     -d '{
           "model": "my-route",
           "messages": [{"role":"user","content":"Say hello"}]
         }'

```

Or use the Anthropic Messages API format:

```bash
curl -X POST http://localhost:4000/v1/messages \
     -H "Content-Type: application/json" \
     -d '{
           "model": "my-anthropic-route",
           "messages": [{"role":"user","content":"Explain quantum entanglement"}]
         }'

```

Inspect routing statistics via the built-in endpoint:

```bash
curl http://localhost:4000/v1/stats | jq .

```

## Summary

- **Server startup** binds TCP sockets and registers Axum routes via `BoundServer::bind` in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), loading configuration from [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs).
- **Request timing** starts with the `stamp_request_start` middleware, which inserts an `Instant` into request extensions for latency tracking.
- **Route resolution** occurs in `resolve_route`, validating models against `ServerState.routes` and decoding JSON via `decode_request` in the translation layer.
- **Algorithm execution** runs through `libsy::drive`, potentially invoking downstream clients recursively via `serve_decision_dependency` and [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs).
- **Response serialization** uses `into_http_response` to generate wire-format-specific JSON and adds the `x-model-router-selected-model` header.
- **Error handling** wraps failures in `ApiError` objects rendered by `render_error_response`, with final logging capturing status codes and durations.

## Frequently Asked Questions

### What Axum middleware does Switchyard use for request timing?

Switchyard uses the `stamp_request_start` middleware located in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs). This middleware inserts a `RequestStart(Instant)` into the request extensions immediately upon ingress, enabling precise latency calculation from TCP acceptance to response serialization.

### How does Switchyard validate incoming model names?

During the `resolve_route` phase in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs), Switchyard validates that the `model` field is non-empty and then performs a lookup via `route_for_model` against the `ServerState.routes` map. If the model is undefined or the authentication kind mismatches the wire format, the request rejects before reaching the routing algorithm.

### Where is the selected model ID exposed in the response?

The `into_http_response` function in [`crates/switchyard-server/src/response.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/response.rs) attaches the selected model ID as the `x-model-router-selected-model` HTTP header. This header appears in the final HTTP response regardless of whether the wire format is OpenAI, Anthropic, or OpenAI-compatible responses.

### How are routing algorithm decisions executed?

The `handle_llm_request` function invokes `switchyard_llm_client::run`, which delegates to `libsy::drive` with the route’s configured algorithm. The algorithm may call intermediate models (classifiers or judges) through `serve_decision_dependency`, which uses [`crates/libsy-llm-client/src/client.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/src/client.rs) to forward requests downstream before returning a final `RoutingOutcome`.