# How to Build, Package, and Configure the NVIDIA NeMo Relay Plugin for Switchyard to Use a Specific routes.toml Deployment

> Learn to build, package, and configure the NVIDIA NeMo Relay plugin for Switchyard. Point your routes.toml deployment and dynamically route LLM requests with this Rust cdylib.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-13

---

**The NeMo Relay plugin is a Rust `cdylib` that loads into the NeMo Relay inference server at runtime, parses a JSON configuration to locate a TOML routing definition, and initializes a `SwitchyardRuntime` to intercept and dynamically route LLM requests.**

The NVIDIA NeMo Relay plugin for Switchyard bridges the NeMo Relay inference server with Switchyard’s intelligent routing engine. Implemented as a dynamic library in the `NVIDIA-NeMo/Switchyard` repository, this plugin converts Relay’s incoming requests into Switchyard routing decisions based on a configurable [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) deployment. Understanding its build process, packaging requirements, and configuration schema is essential for deploying production-grade model routing.

## Building the Plugin as a Rust cdylib

The plugin source resides in `crates/switchyard-nemo-relay-plugin/`. To produce a binary compatible with Relay’s plugin loader, the [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) declares the crate type as a C-compatible dynamic library:

```toml
[lib]
crate-type = ["cdylib"]

```

The `cdylib` setting instructs Cargo to emit a shared object (`.so` on Linux, `.dll` on Windows, `.dylib` on macOS) rather than a static Rust library. This format is required by the Relay plugin system, which uses the `nemo_relay_plugin!` macro to load external routing logic.

Compile the release artifact with:

```bash
cargo build --release -p switchyard-nemo-relay-plugin

```

The output `target/release/libswitchyard-nemo-relay-plugin.so` (or platform equivalent) contains all necessary dependencies—including `futures-util`, `http`, and `toml`— statically linked within the shared library.

## Packaging the Shared Library for Deployment

Because the artifact is a self-contained shared library, packaging involves placing the compiled binary in Relay’s plugin search path. Copy `libswitchyard-nemo-relay-plugin.so` into the directory Relay scans at startup (commonly `relay/plugins/` or the path specified by the `PLUGIN_PATH` environment variable).

For containerized deployments, bundle the `.so` file alongside the Relay server binary in a Docker image. No additional Rust toolchain or source code is required on the target host; the plugin carries its own routing logic and only requires a valid JSON configuration file to locate the Switchyard deployment.

## Configuring the Plugin to Point to a Specific routes.toml Deployment

When the Relay server invokes the plugin’s `register` function in [`src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/lib.rs), the plugin parses a JSON configuration snippet and constructs a `SwitchyardConfig` struct (defined in [`src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/config.rs)). The configuration must specify **exactly one** of two mutually exclusive fields to locate the routing table:

- **`switchyard_config_path`**: A filesystem path to an external TOML file (typically [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml)).
- **`switchyard_config`**: An inline TOML object encoded as JSON for embedded deployments.

The `SwitchyardConfig` struct models these options:

```rust
pub(crate) struct SwitchyardConfig {
    #[serde(default)]
    pub(crate) priority: i32,
    #[serde(default)]
    pub(crate) switchyard_config_path: Option<PathBuf>,
    #[serde(default)]
    pub(crate) switchyard_config: Option<Map<String, Value>>,
}

```

When `SwitchyardRuntime::new` initializes, it calls `config.load_runner()`, which branches based on the provided fields:

```rust
match (&self.switchyard_config_path, &self.switchyard_config) {
    (Some(path), None) => Runner::load(path),
    (None, Some(config)) => {
        let source = toml::to_string(config)?;
        Runner::from_toml(&source)
    }
    // Errors if both or neither are set
}

```

### Loading Routes from an External File with switchyard_config_path

To point the plugin at a concrete [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) on disk, provide a JSON configuration object like:

```json
{
  "priority": 10,
  "switchyard_config_path": "/etc/relay/routes.toml"
}

```

The `Runner::load` function reads the file, validates the schema version, and constructs the routing graph containing `Route`, `Algorithm`, and `ModelCapabilities` definitions.

### Embedding Routes Directly with switchyard_config_path

For environments where file system access is restricted, embed the TOML content as a JSON object:

```json
{
  "priority": 10,
  "switchyard_config": {
    "schema_version": 1,
    "llm_clients": { "primary": { "format": "openai_chat", "base_url": "https://api.openai.com/v1" } },
    "targets": { "default": { "id": "gpt-4", "llm_client": "primary" } },
    "routes": { "default": { "id": "switchyard/default", "type": "passthrough", "target": "default" } }
  }
}

```

The plugin serializes this object to TOML internally and passes it to `Runner::from_toml`.

## Runtime Initialization and Request Interception

After parsing the configuration, the `register` function creates an `Arc<SwitchyardRuntime>` and installs two interceptors:

- **`switchyard.runner.buffered`**: Handles standard request-response LLM calls via `runtime.execute_buffered`.
- **`switchyard.runner.streaming`**: Handles server-sent event (SSE) streams via `runtime.execute_stream`.

Both interceptors follow the same flow defined in [`src/runtime.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/runtime.rs). They first map the Relay call name to a `WireFormat` using `protocol_from_call`, then inspect the request’s `model` field. If `runtime.manages_model(model)` returns true, the interceptor decodes the request with `runtime.decode_request`, executes the routing algorithm, and returns the generated response or stream back to Relay.

## Emitting Telemetry to Relay

All routing decisions, token usage metrics, and error conditions are forwarded to the Relay observability system through `emit_event` and `emit_events` helpers. These functions translate internal `RoutingEvent` structures into Relay-compatible marks and metrics, ensuring complete visibility into Switchyard’s routing behavior from the Relay dashboard.

## Summary

- **Build**: Compile `crates/switchyard-nemo-relay-plugin/` with `crate-type = ["cdylib"]` to generate a shared library for the Relay plugin loader.
- **Package**: Deploy the resulting `.so` file (or platform equivalent) into Relay’s plugin directory; no external Rust dependencies are required at runtime.
- **Configure**: Supply a JSON configuration specifying either `switchyard_config_path` (external TOML file) or `switchyard_config` (inline JSON object), but never both.
- **Initialize**: The plugin’s `register` function parses the JSON, builds a `SwitchyardRuntime`, and loads a `Runner` from the TOML source to create the routing graph.
- **Route**: Interceptors for buffered and streaming calls check `manages_model`, decode requests via the runtime, and execute the appropriate routing logic while emitting telemetry events back to Relay.

## Frequently Asked Questions

### What file format does the NeMo Relay plugin require for Switchyard configuration?

The plugin consumes a **JSON configuration** at the plugin level (supplied to Relay), which must contain either a file path to a **TOML** routing definition or an inline TOML object. The actual routing schema—including `llm_clients`, `targets`, and `routes`—is defined in standard TOML format.

### Can I use both `switchyard_config_path` and `switchyard_config` simultaneously?

No. The `SwitchyardConfig` logic in [`src/config.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/src/config.rs) enforces mutual exclusivity. If both fields are provided, or if neither is provided, the `load_runner()` function returns an error and the plugin fails to initialize.

### Where in the source code does the plugin handle the routing logic for incoming requests?

The entry point is the `register` function in [`crates/switchyard-nemo-relay-plugin/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/lib.rs), which creates the `SwitchyardRuntime`. The actual request processing—decoding, model verification via `manages_model`, and execution—occurs in [`crates/switchyard-nemo-relay-plugin/src/runtime.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/runtime.rs) within the buffered and streaming interceptor implementations.

### How do I rebuild the plugin for a different target architecture?

Use Cargo’s cross-compilation flags with the same build command: `cargo build --release -p switchyard-nemo-relay-plugin --target <target-triple>`. Ensure the target platform supports the `cdylib` crate type, and verify that the Relay server on the target architecture can load the resulting shared library format.