# TOML Configuration File for Switchyard Deployments: Complete Schema Guide

> Understand the TOML configuration file structure for Switchyard deployments. Learn about LLM clients, model targets, and routing endpoints in this comprehensive schema guide.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: api-reference
- Published: 2026-08-22

---

**A Switchyard deployment is controlled by a single TOML file that defines three core concepts: upstream LLM providers in `[llm_clients.<name>]` tables, concrete model instances in `[targets.<name>]` tables, and public routing endpoints in `[routes.<name>]` tables.**

The NVIDIA-NeMo/Switchyard repository implements an LLM routing server that relies entirely on a declarative **TOML configuration file** to orchestrate deployments. This configuration dictates how requests flow from clients through routing algorithms to upstream model providers. Understanding the precise structure of this schema is essential for deploying production-grade Switchyard servers.

## Core Configuration Sections

The Switchyard TOML schema is organized into four distinct top-level sections that work together to define the serving pipeline.

### Schema Version Declaration

Every configuration must begin with `schema_version = 1`. This integer signals the parser version and ensures backward compatibility as the Switchyard server evolves.

```toml
schema_version = 1

```

### LLM Clients Configuration

The **`[llm_clients.<name>]`** tables define upstream provider connections. Each client specifies how Switchyard connects to external APIs like OpenAI, Anthropic, or OpenRouter.

Required fields include:
- **`format`** – The API protocol: `openai_chat`, `openai_responses`, or `anthropic_messages`
- **`base_url`** – The provider's endpoint URL

Authentication uses one of two mutually exclusive methods:
- **`api_key_env`** – Name of an environment variable containing the credential
- **`forward_auth`** – Boolean to forward the caller's original credential

Optional tuning parameters include `extra_headers` for custom HTTP headers and `max_retries` for resilience.

```toml
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
max_retries = 3

```

### Targets Definition

The **`[targets.<name>]`** tables declare concrete model instances that the server can route to. Each target maps to a specific upstream model ID and references a configured LLM client.

Required fields:
- **`id`** – The exact model identifier sent upstream (e.g., `anthropic/claude-sonnet-4.5`)
- **`llm_client`** – A string referencing one of the `[llm_clients]` table names

Optional fields:
- **`extra_body`** – A nested table that merges additional JSON fields into upstream requests, useful for provider-specific parameters like reasoning effort

```toml
[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"

[targets.efficient]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
extra_body = { reasoning = { effort = "medium" } }

```

### Routes and Routing Algorithms

The **`[routes.<name>]`** tables define public model IDs that callers use. Each route selects a routing algorithm via the **`type`** field and provides algorithm-specific configuration.

All routes share common optional keys:
- **`id`** – Public identifier exposed to callers
- **`context_window`** – Token limit for the route
- **`tool_calling`** – Boolean enabling function calling
- **`reasoning`** – Boolean or configuration for reasoning capabilities

Switchyard supports six distinct routing types:

**`noop`** – Returns a static "OK" response for smoke testing infrastructure without consuming LLM tokens.

**`passthrough`** – Forwards every request to a single target. Requires a `target` field referencing a target name, or optionally a nested classifier for sub-agent workflows.

**`random`** – Distributes traffic across multiple targets using weighted random selection. Requires a `targets` array and optional `weights` array and `seed` for reproducibility.

**`llm_classifier`** – Executes a judge model to dynamically select between strong and weak model tiers. Requires:
- **`mode`** – One of `capability`, `escalation`, or `custom`
- **`classifier_target`** – The target used for judgment
- **`strong_target`** and **`weak_target`** – The potential destinations
- **`base_threshold`** – Confidence threshold for routing decisions

**`stage_router`** – Chooses between a capable and efficient tier each turn based on signal confidence. Requires:
- **`capable_target`** and **`efficient_target`**
- **`picker`** – Strategy such as `efficient_first`
- **`confidence_threshold`** – Float value for tier switching
- **`recent_turn_window`** – Integer tracking conversation history

**`advisor`** – Implements a review loop where a weaker executor model generates output and a stronger advisor model requests revisions. Requires:
- **`executor_target`** and **`advisor_target`**
- **`max_reviews`** – Maximum iteration count
- **`gate_stall_turns`** and **`gate_min_tool_results`** – Gatekeeping parameters

## Validation and Testing

Before deploying, validate your TOML structure using the Switchyard server binary. The `--dry-run` flag parses the configuration without starting the service.

```bash
switchyard-server --config <file> --dry-run

```

According to the source code in `crates/switchyard-server`, the configuration must contain at least one `[targets]` table (which may be empty) and one `[routes]` table. The `[llm_clients]` section can be omitted if only local routes are used.

## Configuration Examples

### Minimal Production Configuration

This example from [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) demonstrates a single-provider passthrough setup:

```toml
schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"

[routes.default]
id = "switchyard"
type = "passthrough"
target = "strong"

```

### Full-Featured Deployment

This comprehensive example from [`dev-server/config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/dev-server/config.toml) showcases multiple routing strategies:

```toml
schema_version = 1

[llm_clients.inference_hub]
format = "openai_responses"
base_url = "https://inference-api.nvidia.com/v1"
api_key_env = "NVIDIA_API_KEY"

[targets.capable]
id = "openai/openai/gpt-5.6-sol"
llm_client = "inference_hub"
extra_body = { reasoning = { effort = "medium" } }

[targets.efficient]
id = "openai/openai/gpt-5.6-luna"
llm_client = "inference_hub"

[routes.random]
id = "switchyard/random"
type = "random"
targets = ["capable", "efficient"]

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

[routes.classifier]
id = "switchyard/classifier"
type = "llm_classifier"
mode = "capability"
classifier_target = "capable"
strong_target = "capable"
weak_target = "efficient"
base_threshold = 0.5
classify_trigger = "new_session"

[routes.advisor]
id = "switchyard/advisor"
type = "advisor"
executor_target = "efficient"
advisor_target = "capable"
max_reviews = 3
gate_stall_turns = 30
gate_min_tool_results = 3

```

## Summary

- **Three core sections** define every Switchyard deployment: `[llm_clients]` for upstream connections, `[targets]` for model instances, and `[routes]` for public endpoints.
- **Six routing algorithms** provide capabilities from simple passthrough to intelligent tier selection via `llm_classifier` and `stage_router`.
- **Environment-based authentication** keeps secrets out of version control through the `api_key_env` field.
- **Validation occurs** via `switchyard-server --config <file> --dry-run` before production deployment.
- **Reference implementations** exist in [`dev-server/config.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/dev-server/config.toml) and [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) within the NVIDIA-NeMo/Switchyard repository.

## Frequently Asked Questions

### What is the minimum required TOML configuration for Switchyard?

You must define `schema_version = 1`, at least one `[targets]` table (which may be empty), and at least one `[routes]` table. The `[llm_clients]` section is optional only if you are using local routes that do not require upstream authentication. A valid minimal configuration requires a target with an `id` and `llm_client`, plus a route with `type` and appropriate target references.

### How do I validate a Switchyard TOML configuration file before deploying?

Run the Switchyard server binary with the `--dry-run` flag: `switchyard-server --config <path> --dry-run`. This command, implemented in `crates/switchyard-server`, parses the entire configuration and reports schema errors without starting the HTTP server or consuming API credits.

### Can I use environment variables for API keys in Switchyard TOML files?

Yes. Use the `api_key_env` field within any `[llm_clients.<name>]` table to specify the name of an environment variable containing the API key. Alternatively, set `forward_auth = true` to pass through the caller's original authentication header, which is useful for multi-tenant deployments.

### What is the difference between targets and routes in Switchyard?

**Targets** represent concrete upstream model instances (e.g., `gpt-4o` or `claude-sonnet`) and specify the exact `id` sent to the provider. **Routes** represent public-facing endpoints that callers interact with, defining the routing algorithm (`type`) and selecting which targets to use. A single route can reference multiple targets for load balancing or tiered routing strategies.