TOML Configuration File for Switchyard Deployments: Complete Schema Guide
A Switchyard deployment is controlled by a single TOML file that defines three core concepts: upstream LLM providers in [llm_clients.<name>] tables, concrete model instances in [targets.<name>] tables, and public routing endpoints in [routes.<name>] tables.
The NVIDIA-NeMo/Switchyard repository implements an LLM routing server that relies entirely on a declarative TOML configuration file to orchestrate deployments. This configuration dictates how requests flow from clients through routing algorithms to upstream model providers. Understanding the precise structure of this schema is essential for deploying production-grade Switchyard servers.
Core Configuration Sections
The Switchyard TOML schema is organized into four distinct top-level sections that work together to define the serving pipeline.
Schema Version Declaration
Every configuration must begin with schema_version = 1. This integer signals the parser version and ensures backward compatibility as the Switchyard server evolves.
schema_version = 1
LLM Clients Configuration
The [llm_clients.<name>] tables define upstream provider connections. Each client specifies how Switchyard connects to external APIs like OpenAI, Anthropic, or OpenRouter.
Required fields include:
format– The API protocol:openai_chat,openai_responses, oranthropic_messagesbase_url– The provider's endpoint URL
Authentication uses one of two mutually exclusive methods:
api_key_env– Name of an environment variable containing the credentialforward_auth– Boolean to forward the caller's original credential
Optional tuning parameters include extra_headers for custom HTTP headers and max_retries for resilience.
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
max_retries = 3
Targets Definition
The [targets.<name>] tables declare concrete model instances that the server can route to. Each target maps to a specific upstream model ID and references a configured LLM client.
Required fields:
id– The exact model identifier sent upstream (e.g.,anthropic/claude-sonnet-4.5)llm_client– A string referencing one of the[llm_clients]table names
Optional fields:
extra_body– A nested table that merges additional JSON fields into upstream requests, useful for provider-specific parameters like reasoning effort
[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"
[targets.efficient]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
extra_body = { reasoning = { effort = "medium" } }
Routes and Routing Algorithms
The [routes.<name>] tables define public model IDs that callers use. Each route selects a routing algorithm via the type field and provides algorithm-specific configuration.
All routes share common optional keys:
id– Public identifier exposed to callerscontext_window– Token limit for the routetool_calling– Boolean enabling function callingreasoning– Boolean or configuration for reasoning capabilities
Switchyard supports six distinct routing types:
noop – Returns a static "OK" response for smoke testing infrastructure without consuming LLM tokens.
passthrough – Forwards every request to a single target. Requires a target field referencing a target name, or optionally a nested classifier for sub-agent workflows.
random – Distributes traffic across multiple targets using weighted random selection. Requires a targets array and optional weights array and seed for reproducibility.
llm_classifier – Executes a judge model to dynamically select between strong and weak model tiers. Requires:
mode– One ofcapability,escalation, orcustomclassifier_target– The target used for judgmentstrong_targetandweak_target– The potential destinationsbase_threshold– Confidence threshold for routing decisions
stage_router – Chooses between a capable and efficient tier each turn based on signal confidence. Requires:
capable_targetandefficient_targetpicker– Strategy such asefficient_firstconfidence_threshold– Float value for tier switchingrecent_turn_window– Integer tracking conversation history
advisor – Implements a review loop where a weaker executor model generates output and a stronger advisor model requests revisions. Requires:
executor_targetandadvisor_targetmax_reviews– Maximum iteration countgate_stall_turnsandgate_min_tool_results– Gatekeeping parameters
Validation and Testing
Before deploying, validate your TOML structure using the Switchyard server binary. The --dry-run flag parses the configuration without starting the service.
switchyard-server --config <file> --dry-run
According to the source code in crates/switchyard-server, the configuration must contain at least one [targets] table (which may be empty) and one [routes] table. The [llm_clients] section can be omitted if only local routes are used.
Configuration Examples
Minimal Production Configuration
This example from docs/reference/toml_schema.md demonstrates a single-provider passthrough setup:
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"
[routes.default]
id = "switchyard"
type = "passthrough"
target = "strong"
Full-Featured Deployment
This comprehensive example from dev-server/config.toml showcases multiple routing strategies:
schema_version = 1
[llm_clients.inference_hub]
format = "openai_responses"
base_url = "https://inference-api.nvidia.com/v1"
api_key_env = "NVIDIA_API_KEY"
[targets.capable]
id = "openai/openai/gpt-5.6-sol"
llm_client = "inference_hub"
extra_body = { reasoning = { effort = "medium" } }
[targets.efficient]
id = "openai/openai/gpt-5.6-luna"
llm_client = "inference_hub"
[routes.random]
id = "switchyard/random"
type = "random"
targets = ["capable", "efficient"]
[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3
[routes.classifier]
id = "switchyard/classifier"
type = "llm_classifier"
mode = "capability"
classifier_target = "capable"
strong_target = "capable"
weak_target = "efficient"
base_threshold = 0.5
classify_trigger = "new_session"
[routes.advisor]
id = "switchyard/advisor"
type = "advisor"
executor_target = "efficient"
advisor_target = "capable"
max_reviews = 3
gate_stall_turns = 30
gate_min_tool_results = 3
Summary
- Three core sections define every Switchyard deployment:
[llm_clients]for upstream connections,[targets]for model instances, and[routes]for public endpoints. - Six routing algorithms provide capabilities from simple passthrough to intelligent tier selection via
llm_classifierandstage_router. - Environment-based authentication keeps secrets out of version control through the
api_key_envfield. - Validation occurs via
switchyard-server --config <file> --dry-runbefore production deployment. - Reference implementations exist in
dev-server/config.tomlanddocs/reference/toml_schema.mdwithin the NVIDIA-NeMo/Switchyard repository.
Frequently Asked Questions
What is the minimum required TOML configuration for Switchyard?
You must define schema_version = 1, at least one [targets] table (which may be empty), and at least one [routes] table. The [llm_clients] section is optional only if you are using local routes that do not require upstream authentication. A valid minimal configuration requires a target with an id and llm_client, plus a route with type and appropriate target references.
How do I validate a Switchyard TOML configuration file before deploying?
Run the Switchyard server binary with the --dry-run flag: switchyard-server --config <path> --dry-run. This command, implemented in crates/switchyard-server, parses the entire configuration and reports schema errors without starting the HTTP server or consuming API credits.
Can I use environment variables for API keys in Switchyard TOML files?
Yes. Use the api_key_env field within any [llm_clients.<name>] table to specify the name of an environment variable containing the API key. Alternatively, set forward_auth = true to pass through the caller's original authentication header, which is useful for multi-tenant deployments.
What is the difference between targets and routes in Switchyard?
Targets represent concrete upstream model instances (e.g., gpt-4o or claude-sonnet) and specify the exact id sent to the provider. Routes represent public-facing endpoints that callers interact with, defining the routing algorithm (type) and selecting which targets to use. A single route can reference multiple targets for load balancing or tiered routing strategies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →