TOML Configuration File for Switchyard Deployments: Complete Schema Guide

A Switchyard deployment is controlled by a single TOML file that defines three core concepts: upstream LLM providers in [llm_clients.<name>] tables, concrete model instances in [targets.<name>] tables, and public routing endpoints in [routes.<name>] tables.

The NVIDIA-NeMo/Switchyard repository implements an LLM routing server that relies entirely on a declarative TOML configuration file to orchestrate deployments. This configuration dictates how requests flow from clients through routing algorithms to upstream model providers. Understanding the precise structure of this schema is essential for deploying production-grade Switchyard servers.

Core Configuration Sections

The Switchyard TOML schema is organized into four distinct top-level sections that work together to define the serving pipeline.

Schema Version Declaration

Every configuration must begin with schema_version = 1. This integer signals the parser version and ensures backward compatibility as the Switchyard server evolves.

schema_version = 1

LLM Clients Configuration

The [llm_clients.<name>] tables define upstream provider connections. Each client specifies how Switchyard connects to external APIs like OpenAI, Anthropic, or OpenRouter.

Required fields include:

  • format – The API protocol: openai_chat, openai_responses, or anthropic_messages
  • base_url – The provider's endpoint URL

Authentication uses one of two mutually exclusive methods:

  • api_key_env – Name of an environment variable containing the credential
  • forward_auth – Boolean to forward the caller's original credential

Optional tuning parameters include extra_headers for custom HTTP headers and max_retries for resilience.

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
max_retries = 3

Targets Definition

The [targets.<name>] tables declare concrete model instances that the server can route to. Each target maps to a specific upstream model ID and references a configured LLM client.

Required fields:

  • id – The exact model identifier sent upstream (e.g., anthropic/claude-sonnet-4.5)
  • llm_client – A string referencing one of the [llm_clients] table names

Optional fields:

  • extra_body – A nested table that merges additional JSON fields into upstream requests, useful for provider-specific parameters like reasoning effort
[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"

[targets.efficient]
id = "openai/gpt-4o-mini"
llm_client = "openrouter"
extra_body = { reasoning = { effort = "medium" } }

Routes and Routing Algorithms

The [routes.<name>] tables define public model IDs that callers use. Each route selects a routing algorithm via the type field and provides algorithm-specific configuration.

All routes share common optional keys:

  • id – Public identifier exposed to callers
  • context_window – Token limit for the route
  • tool_calling – Boolean enabling function calling
  • reasoning – Boolean or configuration for reasoning capabilities

Switchyard supports six distinct routing types:

noop – Returns a static "OK" response for smoke testing infrastructure without consuming LLM tokens.

passthrough – Forwards every request to a single target. Requires a target field referencing a target name, or optionally a nested classifier for sub-agent workflows.

random – Distributes traffic across multiple targets using weighted random selection. Requires a targets array and optional weights array and seed for reproducibility.

llm_classifier – Executes a judge model to dynamically select between strong and weak model tiers. Requires:

  • mode – One of capability, escalation, or custom
  • classifier_target – The target used for judgment
  • strong_target and weak_target – The potential destinations
  • base_threshold – Confidence threshold for routing decisions

stage_router – Chooses between a capable and efficient tier each turn based on signal confidence. Requires:

  • capable_target and efficient_target
  • picker – Strategy such as efficient_first
  • confidence_threshold – Float value for tier switching
  • recent_turn_window – Integer tracking conversation history

advisor – Implements a review loop where a weaker executor model generates output and a stronger advisor model requests revisions. Requires:

  • executor_target and advisor_target
  • max_reviews – Maximum iteration count
  • gate_stall_turns and gate_min_tool_results – Gatekeeping parameters

Validation and Testing

Before deploying, validate your TOML structure using the Switchyard server binary. The --dry-run flag parses the configuration without starting the service.

switchyard-server --config <file> --dry-run

According to the source code in crates/switchyard-server, the configuration must contain at least one [targets] table (which may be empty) and one [routes] table. The [llm_clients] section can be omitted if only local routes are used.

Configuration Examples

Minimal Production Configuration

This example from docs/reference/toml_schema.md demonstrates a single-provider passthrough setup:

schema_version = 1

[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"

[targets.strong]
id = "anthropic/claude-sonnet-4.5"
llm_client = "openrouter"

[routes.default]
id = "switchyard"
type = "passthrough"
target = "strong"

This comprehensive example from dev-server/config.toml showcases multiple routing strategies:

schema_version = 1

[llm_clients.inference_hub]
format = "openai_responses"
base_url = "https://inference-api.nvidia.com/v1"
api_key_env = "NVIDIA_API_KEY"

[targets.capable]
id = "openai/openai/gpt-5.6-sol"
llm_client = "inference_hub"
extra_body = { reasoning = { effort = "medium" } }

[targets.efficient]
id = "openai/openai/gpt-5.6-luna"
llm_client = "inference_hub"

[routes.random]
id = "switchyard/random"
type = "random"
targets = ["capable", "efficient"]

[routes.stage]
id = "switchyard/stage"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

[routes.classifier]
id = "switchyard/classifier"
type = "llm_classifier"
mode = "capability"
classifier_target = "capable"
strong_target = "capable"
weak_target = "efficient"
base_threshold = 0.5
classify_trigger = "new_session"

[routes.advisor]
id = "switchyard/advisor"
type = "advisor"
executor_target = "efficient"
advisor_target = "capable"
max_reviews = 3
gate_stall_turns = 30
gate_min_tool_results = 3

Summary

  • Three core sections define every Switchyard deployment: [llm_clients] for upstream connections, [targets] for model instances, and [routes] for public endpoints.
  • Six routing algorithms provide capabilities from simple passthrough to intelligent tier selection via llm_classifier and stage_router.
  • Environment-based authentication keeps secrets out of version control through the api_key_env field.
  • Validation occurs via switchyard-server --config <file> --dry-run before production deployment.
  • Reference implementations exist in dev-server/config.toml and docs/reference/toml_schema.md within the NVIDIA-NeMo/Switchyard repository.

Frequently Asked Questions

What is the minimum required TOML configuration for Switchyard?

You must define schema_version = 1, at least one [targets] table (which may be empty), and at least one [routes] table. The [llm_clients] section is optional only if you are using local routes that do not require upstream authentication. A valid minimal configuration requires a target with an id and llm_client, plus a route with type and appropriate target references.

How do I validate a Switchyard TOML configuration file before deploying?

Run the Switchyard server binary with the --dry-run flag: switchyard-server --config <path> --dry-run. This command, implemented in crates/switchyard-server, parses the entire configuration and reports schema errors without starting the HTTP server or consuming API credits.

Can I use environment variables for API keys in Switchyard TOML files?

Yes. Use the api_key_env field within any [llm_clients.<name>] table to specify the name of an environment variable containing the API key. Alternatively, set forward_auth = true to pass through the caller's original authentication header, which is useful for multi-tenant deployments.

What is the difference between targets and routes in Switchyard?

Targets represent concrete upstream model instances (e.g., gpt-4o or claude-sonnet) and specify the exact id sent to the provider. Routes represent public-facing endpoints that callers interact with, defining the routing algorithm (type) and selecting which targets to use. A single route can reference multiple targets for load balancing or tiered routing strategies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →