# How to Manage Configurations in Switchyard: A Complete TOML-Based Guide

> Master Switchyard configuration management with this complete TOML guide. Learn how Switchyard uses TOML to create typed DeploymentConfig objects for inference and HTTP server.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Switchyard uses a TOML-based configuration system parsed at startup by the `Runner::load` method to build fully-typed `DeploymentConfig` objects that power both the inference library and HTTP server.**

NVIDIA-NeMo/Switchyard drives LLM routing behavior through a strict TOML configuration file that defines clients, targets, and routes. This single file is deserialized into a type-safe `DeploymentConfig` struct according to the source code in [`crates/switchyard-runner/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/config.rs), and validated at startup to ensure all references resolve correctly before any inference calls execute.

## Understanding the Core Configuration Flow

The configuration pipeline follows a strict validation path from file to runtime object, ensuring misconfigurations are caught before serving traffic.

### Loading the TOML Configuration File

The entry point for all Switchyard deployments is `Runner::load`, exposed by the `switchyard-runner` crate. This static method accepts a file path and delegates to `load_runner`, which internally calls `runner_from_toml` to read and parse the file.

```rust
use switchyard_runner::Runner;

let runner = Runner::load("dev-server/config.toml")?;

```

The loader validates the `schema_version` field immediately; current implementations require `schema_version = 1`, and future versions will be rejected during this phase.

### Deserializing into DeploymentConfig

The TOML content deserializes into the `DeploymentConfig` struct defined in [`crates/switchyard-runner/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/config.rs). This struct enforces strict schema compliance through `#[serde(deny_unknown_fields)]`, which rejects any unknown keys and prevents typos or deprecated fields from silently passing through.

The struct contains four primary fields:
- `schema_version`: Integer version marker (must be `1`)
- `llm_clients`: Map of client names to backend configurations
- `targets`: Logical model identifiers referencing specific clients
- `routes`: Routing definitions connecting requests to targets

### Validating and Building Internal Objects

After deserialization, the `build` method (implemented around lines 55–150 in [`config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/config.rs)) walks the configuration graph to construct runtime objects:

- **Client construction**: Calls `build_clients` to create `TranslatingLlmClient` instances for each LLM backend defined in `llm_clients`
- **Target resolution**: Executes `build_targets` to map target names to model IDs, verifying each target references an existing LLM client
- **Route assembly**: Validates routing algorithms (such as `random`, `stage_router`, or `passthrough`) and builds client routers via `build_route_clients`
- **Error handling**: Wraps all validation failures in `RunnerError::configuration` with descriptive messages indicating exactly which section failed

The final output is a fully initialized `Runner` containing a list of `Route` objects and an optional fallback URL, ready for inference calls.

## Configuration Schema Reference

The TOML file organizes deployment topology into four distinct sections:

**`llm_clients`** defines backend communication parameters:
- `format`: Protocol variant (e.g., `"openai_responses"`, `"openai_chat"`)
- `base_url`: Endpoint for the LLM API
- `api_key_env`: Optional environment variable name containing credentials
- `forward_auth`: Boolean enabling credential forwarding from the caller

**`targets`** assigns logical names to specific models:
- `id`: Full model identifier (e.g., `"openai/openai/gpt-5.6-sol"`)
- `llm_client`: Reference to a key in the `llm_clients` table
- `extra_body`: Optional parameters like reasoning effort levels

**`routes`** determines traffic flow:
- `id`: Unique route identifier
- `type`: Routing algorithm (`random`, `stage_router`, `passthrough`)
- `target` or `targets`: References to entries in the `targets` table
- `capable_target`: Specific field for stage routers indicating the high-capability backend

**`schema_version`** must be set to `1` for all current deployments.

## Server-Side Configuration Loading

The HTTP server binary (`switchyard-server`) consumes the same TOML file through a thin compatibility layer. In [`crates/switchyard-server/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/config.rs), the `load_server_state` function wraps `Runner::load` and adapts the result into a `ServerState`:

```rust
pub fn load_server_state(path: impl AsRef<Path>) -> ServerResult<ServerState> {
    ServerState::from_runner(Runner::load(path).map_err(ServerError::from)?)
}

```

This design ensures that running the server requires only specifying the configuration path:

```bash
switchyard-server --config path/to/custom.toml

```

Whether embedding Switchyard as a library or running the standalone server, the same validation logic and TOML format apply.

## Advanced Configuration Patterns

Beyond file-based loading, Switchyard supports embedded configurations and specialized routing behaviors.

### Inline Configuration for Python

The `switchyard-nemo-relay-plugin` enables embedding TOML directly in Python code without maintaining separate files. The `SwitchyardConfig` class accepts either a file path or an inline configuration string, internally calling `Runner::from_toml`:

```python
from switchyard_nemo_relay_plugin import SwitchyardConfig

cfg = SwitchyardConfig(
    switchyard_config="""
schema_version = 1
[llm_clients.primary]
format = "openai_chat"
base_url = "https://example.test/v1"

[targets.default]
id = "example/model"
llm_client = "primary"

[routes.default]
id = "switchyard/default"
type = "passthrough"
target = "default"
"""
)
runner = cfg.load_runner()

```

### Fallback URL Handling

When a request cannot be satisfied by any defined route, the `Runner` invokes `fallback_base_url()` to determine where to forward the traffic. This mechanism creates a "pass-through" chain to external LLM providers when no local targets match the request requirements.

### Authentication Forwarding

Setting `forward_auth = true` on an LLM client configuration causes the server to proxy the caller's OpenAI or Anthropic credentials to the backend. The `build_route_clients` validation logic ensures that only one credential type is forwarded per route, preventing ambiguous authentication states.

## Summary

- Switchyard uses a strictly-validated **TOML configuration file** with `schema_version = 1` to define all deployment parameters
- The **`Runner::load`** method in [`crates/switchyard-runner/src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/config.rs) parses and validates the file, while **`load_server_state`** provides a server-side wrapper
- **Unknown fields are rejected** via `#[serde(deny_unknown_fields)]`, catching configuration errors at startup rather than runtime
- **Python embedding** is supported through the NEMO relay plugin using `Runner::from_toml`
- **Authentication forwarding** and **fallback URLs** provide production-ready routing flexibility without requiring code changes

## Frequently Asked Questions

### What file format does Switchyard use for configuration?

Switchyard uses **TOML** (Tom's Obvious, Minimal Language) for all configuration files. The system expects a mandatory `schema_version` field set to `1`, along with tables defining `llm_clients`, `targets`, and `routes`. The strict schema validation rejects unknown fields to prevent configuration drift.

### How does Switchyard validate configuration at startup?

Validation occurs in three phases: first, the TOML parser checks syntax and deserializes into the `DeploymentConfig` struct with `deny_unknown_fields` enabled; second, the `build` method verifies that all targets reference valid LLM clients and constructs `TranslatingLlmClient` objects; third, route construction validates that routing algorithms have required parameters. Any failure returns a `RunnerError::configuration` with a descriptive message.

### Can I embed Switchyard configuration directly in Python code?

Yes. The `switchyard-nemo-relay-plugin` crate provides a `SwitchyardConfig` class that accepts an inline TOML string through the `switchyard_config` parameter. This snippet is passed directly to `Runner::from_toml`, bypassing the file system entirely while maintaining the same validation guarantees as file-based loading.

### What happens if a route references a non-existent target?

The `build` method in [`config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/config.rs) explicitly checks that every route's target references exist in the `targets` map. If a route points to an undefined target, the configuration fails to load with a `RunnerError::configuration` indicating the missing reference, preventing the `Runner` from starting in an inconsistent state.