How to Deploy and Use a Self-Hosted OpenAI-Compatible Model Server (vLLM) with Switchyard

Switchyard acts as a thin routing layer that forwards OpenAI-compatible requests to self-hosted backends like vLLM through simple TOML configuration, requiring no code changes to integrate local model servers.

Switchyard is a thin routing layer designed to sit between client applications and LLM backends that expose an OpenAI-compatible API. Because the routing logic only cares about HTTP endpoints and OpenAI-style request/response schemas, you can point Switchyard at a locally running model server such as vLLM without any special integration code. This guide walks through the complete setup using the NVIDIA-NeMo/Switchyard repository, referencing actual source paths and configuration files.

Understanding the Three-Layer Routing Architecture

Switchyard uses three distinct configuration layers to route requests, as documented in docs/routing_algorithms/overview.md. Understanding these layers is essential for connecting a self-hosted vLLM instance.

LLM Clients

An LLM client defines how Switchyard communicates with a provider. For a local vLLM server, you declare a client with format = "openai_chat" and point base_url to your vLLM endpoint (typically http://localhost:8000/v1). Unlike cloud providers, no api_key_env is required because local servers usually do not enforce authentication.

Targets

A target ties a specific model identifier to an LLM client. When you create a target, you assign it an id (such as my-rl-qwen) and reference the client you defined earlier. This abstraction allows you to swap backends without changing your route definitions.

Routes

A route maps a public model ID (the one users send in the model= parameter) to a target. For self-hosted deployments, you typically use a passthrough route type that forwards requests unchanged to your vLLM instance.

Launching vLLM with Hidden-State Extraction

To demonstrate advanced integration, this section shows how to launch vLLM with hidden-state extraction enabled. This configuration is documented in docs/vllm-serve-hidden-state.md and requires specific flags for speculative decoding and KV transfer.

Docker Deployment

Run vLLM in a container with GPU access and mount directories for HuggingFace cache and hidden states storage:

export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
export HF_CACHE_DIR=/tmp/vllm-hf-cache
mkdir -p "${HIDDEN_STATES_DIR}" "${HF_CACHE_DIR}"

docker run -d --name vllm_qwen35 \
  --gpus all \
  --ipc=host \
  -p 0.0.0.0:8000:8000 \
  -v "${HF_CACHE_DIR}:/root/.cache/huggingface" \
  -v "${HIDDEN_STATES_DIR}:${HIDDEN_STATES_DIR}" \
  vllm/vllm-openai:latest-cu129 \
  Qwen/Qwen3.6-35B-A3B \
  --tensor-parallel-size 8 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --no-enable-chunked-prefill \
  --speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
  --kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'

Native CLI Deployment

If you have a compatible vLLM build installed locally, run the server directly:

export HIDDEN_STATES_DIR=/tmp/vllm-hidden-states
mkdir -p "${HIDDEN_STATES_DIR}"

vllm serve Qwen/Qwen3.6-35B-A3B \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 32768 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --no-enable-chunked-prefill \
  --speculative-config '{"method":"extract_hidden_states","num_speculative_tokens":1,"draft_model_config":{"hf_config":{"eagle_aux_hidden_state_layer_ids":[39]}}}' \
  --kv-transfer-config '{"kv_connector":"ExampleHiddenStatesConnector","kv_role":"kv_producer","kv_connector_extra_config":{"shared_storage_path":"/tmp/vllm-hidden-states"}}'

Configuring Switchyard for Local vLLM

Create a routes.toml file in your project root to wire the local vLLM instance into Switchyard's routing layer. The complete schema reference is available in docs/reference/toml_schema.md.

Defining the LLM Client

Add a client section that points to your vLLM endpoint:

[llm_clients.local_vllm]
format = "openai_chat"
base_url = "http://localhost:8000/v1"

Creating Targets and Routes

Link the client to a target and expose it via a passthrough route:

[targets.local]
id = "my-rl-qwen"
llm_client = "local_vllm"

[routes.local]
id = "local"
type = "passthrough"
target = "local"

Switchyard validates this configuration at startup and will refuse to start if required fields are missing.

Starting the Switchyard Server

The server binary is built from the switchyard-server crate, with the entry point in crates/switchyard-server/src/main.rs.

Configuration Validation

Validate your TOML syntax and configuration before starting the server:

switchyard-server --config routes.toml --dry-run

Server Startup

Launch the server on port 4000 (or your preferred port):

switchyard-server --config routes.toml --host 127.0.0.1 --port 4000

The server now accepts OpenAI-compatible requests on port 4000 and forwards them to your vLLM instance on port 8000.

Testing the Deployment

Send a chat completion request through Switchyard to verify the routing works:

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model":"local",
    "messages":[{"role":"user","content":"Explain the capital of France."}]
  }'

If you enabled hidden-state extraction in vLLM, the response will contain a kv_transfer_params.hidden_states_path field pointing to a .safetensors file under /tmp/vllm-hidden-states.

Summary

Switchyard simplifies self-hosted LLM deployment by treating vLLM as just another OpenAI-compatible endpoint.

  • Three-layer architecture: Configure an LLM client for the vLLM endpoint, a target to name the model, and a route to expose it publicly.
  • Zero code changes: Integration requires only TOML entries in routes.toml referencing http://localhost:8000/v1.
  • Full feature support: Because Switchyard forwards requests unchanged, vLLM-specific features like hidden-state extraction via --speculative-config and --kv-transfer-config work automatically.
  • Validation: Use --dry-run to validate configurations before starting the server.

Frequently Asked Questions

Do I need to modify Switchyard source code to use local vLLM?

No. As implemented in NVIDIA-NeMo/Switchyard, the routing layer is agnostic to the backend implementation. You only need to add a client definition in routes.toml with format = "openai_chat" and the appropriate base_url. No changes to crates/switchyard-server/src/main.rs or any other source file are required.

How does Switchyard handle vLLM-specific features like hidden-state extraction?

Switchyard forwards the JSON request and response payload unchanged to the backend. When vLLM is launched with --speculative-config and --kv-transfer-config flags, it returns additional metadata such as kv_transfer_params.hidden_states_path in the response. Switchyard passes this through transparently because it implements a passthrough routing strategy that does not modify the OpenAI-compatible schema.

What port does Switchyard listen on by default?

The switchyard-server binary accepts a --port argument to specify the listening port. In the examples from docs/routing_algorithms/overview.md, the server binds to port 4000 using --host 127.0.0.1 --port 4000, while the vLLM backend typically runs on port 8000.

Can I use Switchyard with multiple local vLLM instances?

Yes. Define separate LLM clients for each vLLM instance (differentiating by base_url if on different ports), create distinct targets for each client, and assign unique route IDs. Switchyard will route requests based on the model parameter in the incoming request, allowing you to load balance or segment traffic across multiple self-hosted model servers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →