How to Manage Configurations in Switchyard: A Complete TOML-Based Guide

Switchyard uses a TOML-based configuration system parsed at startup by the Runner::load method to build fully-typed DeploymentConfig objects that power both the inference library and HTTP server.

NVIDIA-NeMo/Switchyard drives LLM routing behavior through a strict TOML configuration file that defines clients, targets, and routes. This single file is deserialized into a type-safe DeploymentConfig struct according to the source code in crates/switchyard-runner/src/config.rs, and validated at startup to ensure all references resolve correctly before any inference calls execute.

Understanding the Core Configuration Flow

The configuration pipeline follows a strict validation path from file to runtime object, ensuring misconfigurations are caught before serving traffic.

Loading the TOML Configuration File

The entry point for all Switchyard deployments is Runner::load, exposed by the switchyard-runner crate. This static method accepts a file path and delegates to load_runner, which internally calls runner_from_toml to read and parse the file.

use switchyard_runner::Runner;

let runner = Runner::load("dev-server/config.toml")?;

The loader validates the schema_version field immediately; current implementations require schema_version = 1, and future versions will be rejected during this phase.

Deserializing into DeploymentConfig

The TOML content deserializes into the DeploymentConfig struct defined in crates/switchyard-runner/src/config.rs. This struct enforces strict schema compliance through #[serde(deny_unknown_fields)], which rejects any unknown keys and prevents typos or deprecated fields from silently passing through.

The struct contains four primary fields:

  • schema_version: Integer version marker (must be 1)
  • llm_clients: Map of client names to backend configurations
  • targets: Logical model identifiers referencing specific clients
  • routes: Routing definitions connecting requests to targets

Validating and Building Internal Objects

After deserialization, the build method (implemented around lines 55–150 in config.rs) walks the configuration graph to construct runtime objects:

  • Client construction: Calls build_clients to create TranslatingLlmClient instances for each LLM backend defined in llm_clients
  • Target resolution: Executes build_targets to map target names to model IDs, verifying each target references an existing LLM client
  • Route assembly: Validates routing algorithms (such as random, stage_router, or passthrough) and builds client routers via build_route_clients
  • Error handling: Wraps all validation failures in RunnerError::configuration with descriptive messages indicating exactly which section failed

The final output is a fully initialized Runner containing a list of Route objects and an optional fallback URL, ready for inference calls.

Configuration Schema Reference

The TOML file organizes deployment topology into four distinct sections:

llm_clients defines backend communication parameters:

  • format: Protocol variant (e.g., "openai_responses", "openai_chat")
  • base_url: Endpoint for the LLM API
  • api_key_env: Optional environment variable name containing credentials
  • forward_auth: Boolean enabling credential forwarding from the caller

targets assigns logical names to specific models:

  • id: Full model identifier (e.g., "openai/openai/gpt-5.6-sol")
  • llm_client: Reference to a key in the llm_clients table
  • extra_body: Optional parameters like reasoning effort levels

routes determines traffic flow:

  • id: Unique route identifier
  • type: Routing algorithm (random, stage_router, passthrough)
  • target or targets: References to entries in the targets table
  • capable_target: Specific field for stage routers indicating the high-capability backend

schema_version must be set to 1 for all current deployments.

Server-Side Configuration Loading

The HTTP server binary (switchyard-server) consumes the same TOML file through a thin compatibility layer. In crates/switchyard-server/src/config.rs, the load_server_state function wraps Runner::load and adapts the result into a ServerState:

pub fn load_server_state(path: impl AsRef<Path>) -> ServerResult<ServerState> {
    ServerState::from_runner(Runner::load(path).map_err(ServerError::from)?)
}

This design ensures that running the server requires only specifying the configuration path:

switchyard-server --config path/to/custom.toml

Whether embedding Switchyard as a library or running the standalone server, the same validation logic and TOML format apply.

Advanced Configuration Patterns

Beyond file-based loading, Switchyard supports embedded configurations and specialized routing behaviors.

Inline Configuration for Python

The switchyard-nemo-relay-plugin enables embedding TOML directly in Python code without maintaining separate files. The SwitchyardConfig class accepts either a file path or an inline configuration string, internally calling Runner::from_toml:

from switchyard_nemo_relay_plugin import SwitchyardConfig

cfg = SwitchyardConfig(
    switchyard_config="""
schema_version = 1
[llm_clients.primary]
format = "openai_chat"
base_url = "https://example.test/v1"

[targets.default]
id = "example/model"
llm_client = "primary"

[routes.default]
id = "switchyard/default"
type = "passthrough"
target = "default"
"""
)
runner = cfg.load_runner()

Fallback URL Handling

When a request cannot be satisfied by any defined route, the Runner invokes fallback_base_url() to determine where to forward the traffic. This mechanism creates a "pass-through" chain to external LLM providers when no local targets match the request requirements.

Authentication Forwarding

Setting forward_auth = true on an LLM client configuration causes the server to proxy the caller's OpenAI or Anthropic credentials to the backend. The build_route_clients validation logic ensures that only one credential type is forwarded per route, preventing ambiguous authentication states.

Summary

  • Switchyard uses a strictly-validated TOML configuration file with schema_version = 1 to define all deployment parameters
  • The Runner::load method in crates/switchyard-runner/src/config.rs parses and validates the file, while load_server_state provides a server-side wrapper
  • Unknown fields are rejected via #[serde(deny_unknown_fields)], catching configuration errors at startup rather than runtime
  • Python embedding is supported through the NEMO relay plugin using Runner::from_toml
  • Authentication forwarding and fallback URLs provide production-ready routing flexibility without requiring code changes

Frequently Asked Questions

What file format does Switchyard use for configuration?

Switchyard uses TOML (Tom's Obvious, Minimal Language) for all configuration files. The system expects a mandatory schema_version field set to 1, along with tables defining llm_clients, targets, and routes. The strict schema validation rejects unknown fields to prevent configuration drift.

How does Switchyard validate configuration at startup?

Validation occurs in three phases: first, the TOML parser checks syntax and deserializes into the DeploymentConfig struct with deny_unknown_fields enabled; second, the build method verifies that all targets reference valid LLM clients and constructs TranslatingLlmClient objects; third, route construction validates that routing algorithms have required parameters. Any failure returns a RunnerError::configuration with a descriptive message.

Can I embed Switchyard configuration directly in Python code?

Yes. The switchyard-nemo-relay-plugin crate provides a SwitchyardConfig class that accepts an inline TOML string through the switchyard_config parameter. This snippet is passed directly to Runner::from_toml, bypassing the file system entirely while maintaining the same validation guarantees as file-based loading.

What happens if a route references a non-existent target?

The build method in config.rs explicitly checks that every route's target references exist in the targets map. If a route points to an undefined target, the configuration fails to load with a RunnerError::configuration indicating the missing reference, preventing the Runner from starting in an inconsistent state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →