# How to Deploy and Use a Self-Hosted OpenAI-Compatible Model Server (vLLM) with Switchyard

> Easily deploy and use a self-hosted OpenAI-compatible model server like vLLM with Switchyard. This routing layer requires no code changes for seamless integration of local models via simple TOML configuration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-22

---

**Switchyard acts as a thin routing layer that forwards OpenAI-compatible requests to self-hosted backends like vLLM through simple TOML configuration, requiring no code changes to integrate local model servers.**

Switchyard is a thin routing layer designed to sit between client applications and LLM backends that expose an OpenAI-compatible API. Because the routing logic only cares about HTTP endpoints and OpenAI-style request/response schemas, you can point Switchyard at a locally running model server such as vLLM without any special integration code. This guide walks through the complete setup using the NVIDIA-NeMo/Switchyard repository, referencing actual source paths and configuration files.

## Understanding the Three-Layer Routing Architecture

Switchyard uses three distinct configuration layers to route requests, as documented in [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md). Understanding these layers is essential for connecting a self-hosted vLLM instance.

### LLM Clients

An **LLM client** defines how Switchyard communicates with a provider. For a local vLLM server, you declare a client with `format = "openai_chat"` and point `base_url` to your vLLM endpoint (typically `http://localhost:8000/v1`). Unlike cloud providers, no `api_key_env` is required because local servers usually do not enforce authentication.

### Targets

A **target** ties a specific model identifier to an LLM client. When you create a target, you assign it an `id` (such as `my-rl-qwen`) and reference the client you defined earlier. This abstraction allows you to swap backends without changing your route definitions.

### Routes

A **route** maps a public model ID (the one users send in the `model=` parameter) to a target. For self-hosted deployments, you typically use a `passthrough` route type that forwards requests unchanged to your vLLM instance.

## Launching vLLM with Hidden-State Extraction

To demonstrate advanced integration, this section shows how to launch vLLM with hidden-state extraction enabled. This configuration is documented in [`docs/vllm-serve-hidden-state.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/vllm-serve-hidden-state.md) and requires specific flags for speculative decoding and KV transfer.

### Docker Deployment

Run vLLM in a container with GPU access and mount directories for HuggingFace cache and hidden states storage:

```bash
export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
export HF_CACHE_DIR=/tmp/vllm-hf-cache
mkdir -p "${HIDDEN_STATES_DIR}" "${HF_CACHE_DIR}"

docker run -d --name vllm_qwen35 \
  --gpus all \
  --ipc=host \
  -p 0.0.0.0:8000:8000 \
  -v "${HF_CACHE_DIR}:/root/.cache/huggingface" \
  -v "${HIDDEN_STATES_DIR}:${HIDDEN_STATES_DIR}" \
  vllm/vllm-openai:latest-cu129 \
  Qwen/Qwen3.6-35B-A3B \
  --tensor-parallel-size 8 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --no-enable-chunked-prefill \
  --speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
  --kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'

```

### Native CLI Deployment

If you have a compatible vLLM build installed locally, run the server directly:

```bash
export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
mkdir -p "${HIDDEN_STATES_DIR}"

vllm serve Qwen/Qwen3.6-35B-A3B \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --no-enable-chunked-prefill \
  --speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
  --kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'

```

## Configuring Switchyard for Local vLLM

Create a [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) file in your project root to wire the local vLLM instance into Switchyard's routing layer. The complete schema reference is available in [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md).

### Defining the LLM Client

Add a client section that points to your vLLM endpoint:

```toml
[llm_clients.local_vllm]
format = "openai_chat"
base_url = "http://localhost:8000/v1"

```

### Creating Targets and Routes

Link the client to a target and expose it via a passthrough route:

```toml
[targets.local]
id = "my-rl-qwen"
llm_client = "local_vllm"

[routes.local]
id = "local"
type = "passthrough"
target = "local"

```

Switchyard validates this configuration at startup and will refuse to start if required fields are missing.

## Starting the Switchyard Server

The server binary is built from the `switchyard-server` crate, with the entry point in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs).

### Configuration Validation

Validate your TOML syntax and configuration before starting the server:

```bash
switchyard-server --config routes.toml --dry-run

```

### Server Startup

Launch the server on port 4000 (or your preferred port):

```bash
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

```

The server now accepts OpenAI-compatible requests on port 4000 and forwards them to your vLLM instance on port 8000.

## Testing the Deployment

Send a chat completion request through Switchyard to verify the routing works:

```bash
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model":"local",
    "messages":[{"role":"user","content":"Explain the capital of France."}]
  }'

```

If you enabled hidden-state extraction in vLLM, the response will contain a `kv_transfer_params.hidden_states_path` field pointing to a `.safetensors` file under `/tmp/vllm-hidden-states`.

## Summary

Switchyard simplifies self-hosted LLM deployment by treating vLLM as just another OpenAI-compatible endpoint.

- **Three-layer architecture**: Configure an **LLM client** for the vLLM endpoint, a **target** to name the model, and a **route** to expose it publicly.
- **Zero code changes**: Integration requires only TOML entries in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) referencing `http://localhost:8000/v1`.
- **Full feature support**: Because Switchyard forwards requests unchanged, vLLM-specific features like hidden-state extraction via `--speculative-config` and `--kv-transfer-config` work automatically.
- **Validation**: Use `--dry-run` to validate configurations before starting the server.

## Frequently Asked Questions

### Do I need to modify Switchyard source code to use local vLLM?

No. As implemented in NVIDIA-NeMo/Switchyard, the routing layer is agnostic to the backend implementation. You only need to add a client definition in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) with `format = "openai_chat"` and the appropriate `base_url`. No changes to [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs) or any other source file are required.

### How does Switchyard handle vLLM-specific features like hidden-state extraction?

Switchyard forwards the JSON request and response payload unchanged to the backend. When vLLM is launched with `--speculative-config` and `--kv-transfer-config` flags, it returns additional metadata such as `kv_transfer_params.hidden_states_path` in the response. Switchyard passes this through transparently because it implements a **passthrough** routing strategy that does not modify the OpenAI-compatible schema.

### What port does Switchyard listen on by default?

The `switchyard-server` binary accepts a `--port` argument to specify the listening port. In the examples from [`docs/routing_algorithms/overview.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/overview.md), the server binds to port 4000 using `--host 127.0.0.1 --port 4000`, while the vLLM backend typically runs on port 8000.

### Can I use Switchyard with multiple local vLLM instances?

Yes. Define separate **LLM clients** for each vLLM instance (differentiating by `base_url` if on different ports), create distinct **targets** for each client, and assign unique route IDs. Switchyard will route requests based on the `model` parameter in the incoming request, allowing you to load balance or segment traffic across multiple self-hosted model servers.