How to Manage Configurations in Switchyard: A Complete TOML-Based Guide
Switchyard uses a TOML-based configuration system parsed at startup by the Runner::load method to build fully-typed DeploymentConfig objects that power both the inference library and HTTP server.
NVIDIA-NeMo/Switchyard drives LLM routing behavior through a strict TOML configuration file that defines clients, targets, and routes. This single file is deserialized into a type-safe DeploymentConfig struct according to the source code in crates/switchyard-runner/src/config.rs, and validated at startup to ensure all references resolve correctly before any inference calls execute.
Understanding the Core Configuration Flow
The configuration pipeline follows a strict validation path from file to runtime object, ensuring misconfigurations are caught before serving traffic.
Loading the TOML Configuration File
The entry point for all Switchyard deployments is Runner::load, exposed by the switchyard-runner crate. This static method accepts a file path and delegates to load_runner, which internally calls runner_from_toml to read and parse the file.
use switchyard_runner::Runner;
let runner = Runner::load("dev-server/config.toml")?;
The loader validates the schema_version field immediately; current implementations require schema_version = 1, and future versions will be rejected during this phase.
Deserializing into DeploymentConfig
The TOML content deserializes into the DeploymentConfig struct defined in crates/switchyard-runner/src/config.rs. This struct enforces strict schema compliance through #[serde(deny_unknown_fields)], which rejects any unknown keys and prevents typos or deprecated fields from silently passing through.
The struct contains four primary fields:
schema_version: Integer version marker (must be1)llm_clients: Map of client names to backend configurationstargets: Logical model identifiers referencing specific clientsroutes: Routing definitions connecting requests to targets
Validating and Building Internal Objects
After deserialization, the build method (implemented around lines 55–150 in config.rs) walks the configuration graph to construct runtime objects:
- Client construction: Calls
build_clientsto createTranslatingLlmClientinstances for each LLM backend defined inllm_clients - Target resolution: Executes
build_targetsto map target names to model IDs, verifying each target references an existing LLM client - Route assembly: Validates routing algorithms (such as
random,stage_router, orpassthrough) and builds client routers viabuild_route_clients - Error handling: Wraps all validation failures in
RunnerError::configurationwith descriptive messages indicating exactly which section failed
The final output is a fully initialized Runner containing a list of Route objects and an optional fallback URL, ready for inference calls.
Configuration Schema Reference
The TOML file organizes deployment topology into four distinct sections:
llm_clients defines backend communication parameters:
format: Protocol variant (e.g.,"openai_responses","openai_chat")base_url: Endpoint for the LLM APIapi_key_env: Optional environment variable name containing credentialsforward_auth: Boolean enabling credential forwarding from the caller
targets assigns logical names to specific models:
id: Full model identifier (e.g.,"openai/openai/gpt-5.6-sol")llm_client: Reference to a key in thellm_clientstableextra_body: Optional parameters like reasoning effort levels
routes determines traffic flow:
id: Unique route identifiertype: Routing algorithm (random,stage_router,passthrough)targetortargets: References to entries in thetargetstablecapable_target: Specific field for stage routers indicating the high-capability backend
schema_version must be set to 1 for all current deployments.
Server-Side Configuration Loading
The HTTP server binary (switchyard-server) consumes the same TOML file through a thin compatibility layer. In crates/switchyard-server/src/config.rs, the load_server_state function wraps Runner::load and adapts the result into a ServerState:
pub fn load_server_state(path: impl AsRef<Path>) -> ServerResult<ServerState> {
ServerState::from_runner(Runner::load(path).map_err(ServerError::from)?)
}
This design ensures that running the server requires only specifying the configuration path:
switchyard-server --config path/to/custom.toml
Whether embedding Switchyard as a library or running the standalone server, the same validation logic and TOML format apply.
Advanced Configuration Patterns
Beyond file-based loading, Switchyard supports embedded configurations and specialized routing behaviors.
Inline Configuration for Python
The switchyard-nemo-relay-plugin enables embedding TOML directly in Python code without maintaining separate files. The SwitchyardConfig class accepts either a file path or an inline configuration string, internally calling Runner::from_toml:
from switchyard_nemo_relay_plugin import SwitchyardConfig
cfg = SwitchyardConfig(
switchyard_config="""
schema_version = 1
[llm_clients.primary]
format = "openai_chat"
base_url = "https://example.test/v1"
[targets.default]
id = "example/model"
llm_client = "primary"
[routes.default]
id = "switchyard/default"
type = "passthrough"
target = "default"
"""
)
runner = cfg.load_runner()
Fallback URL Handling
When a request cannot be satisfied by any defined route, the Runner invokes fallback_base_url() to determine where to forward the traffic. This mechanism creates a "pass-through" chain to external LLM providers when no local targets match the request requirements.
Authentication Forwarding
Setting forward_auth = true on an LLM client configuration causes the server to proxy the caller's OpenAI or Anthropic credentials to the backend. The build_route_clients validation logic ensures that only one credential type is forwarded per route, preventing ambiguous authentication states.
Summary
- Switchyard uses a strictly-validated TOML configuration file with
schema_version = 1to define all deployment parameters - The
Runner::loadmethod incrates/switchyard-runner/src/config.rsparses and validates the file, whileload_server_stateprovides a server-side wrapper - Unknown fields are rejected via
#[serde(deny_unknown_fields)], catching configuration errors at startup rather than runtime - Python embedding is supported through the NEMO relay plugin using
Runner::from_toml - Authentication forwarding and fallback URLs provide production-ready routing flexibility without requiring code changes
Frequently Asked Questions
What file format does Switchyard use for configuration?
Switchyard uses TOML (Tom's Obvious, Minimal Language) for all configuration files. The system expects a mandatory schema_version field set to 1, along with tables defining llm_clients, targets, and routes. The strict schema validation rejects unknown fields to prevent configuration drift.
How does Switchyard validate configuration at startup?
Validation occurs in three phases: first, the TOML parser checks syntax and deserializes into the DeploymentConfig struct with deny_unknown_fields enabled; second, the build method verifies that all targets reference valid LLM clients and constructs TranslatingLlmClient objects; third, route construction validates that routing algorithms have required parameters. Any failure returns a RunnerError::configuration with a descriptive message.
Can I embed Switchyard configuration directly in Python code?
Yes. The switchyard-nemo-relay-plugin crate provides a SwitchyardConfig class that accepts an inline TOML string through the switchyard_config parameter. This snippet is passed directly to Runner::from_toml, bypassing the file system entirely while maintaining the same validation guarantees as file-based loading.
What happens if a route references a non-existent target?
The build method in config.rs explicitly checks that every route's target references exist in the targets map. If a route points to an undefined target, the configuration fails to load with a RunnerError::configuration indicating the missing reference, preventing the Runner from starting in an inconsistent state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →