# How to Fine-Tune Models with Switchyard: A Complete Integration Guide

> Learn how to fine-tune models with Switchyard. Integrate custom models by deploying checkpoints and routing traffic with this comprehensive guide. Get started today!

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**To fine-tune models with Switchyard, train externally using NVIDIA NeMo or similar frameworks, deploy the checkpoint behind an OpenAI-compatible endpoint, register it as a target in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml), and route traffic using Switchyard's algorithms.**

Switchyard is an LLM routing library—not a training framework—that intelligently distributes requests across multiple language models. According to the NVIDIA-NeMo/Switchyard source code, you complete fine-tuning workflows outside the system, then integrate the deployed models into Switchyard's configuration-driven routing layer. This architecture decouples model development from production traffic management, allowing you to swap fine-tuned checkpoints without modifying application code.

## Understanding Switchyard's Architecture

### Switchyard Is a Router, Not a Trainer

Switchyard functions exclusively as a request routing engine. As implemented in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs), the runtime parses TOML configurations and drives algorithms like `stage_router`, but performs no backpropagation or weight updates. The system expects all models—including fine-tuned variants—to be served via external HTTP endpoints that expose LLM-compatible APIs.

## Step 1: Fine-Tune Your Model Using NVIDIA NeMo

Since Switchyard does not provide training capabilities, use NVIDIA NeMo or any ML framework capable of exporting to a servable format. The following example demonstrates supervised fine-tuning using NeMo's prompt learning model:

```python

# Example using NVIDIA‑NeMo (requires nemo‑core & nemo‑model‑zoo)

from nemo.collections.nlp.models.language_modeling import GPTPromptLearningModel

model = GPTPromptLearningModel.restore_from(
    "base_model.nemo"
).to("cuda")

# Prepare your SFT dataset (json / jsonl with {"prompt":..., "completion":...})

model.train_dataset = "sft_dataset.jsonl"

trainer = nemo.utils.Trainer(
    max_epochs=3,
    gpus=1,
    precision=16,
)
trainer.fit(model)

# Export the fine‑tuned checkpoint

model.save_to("fine_tuned_model.nemo")

```

This training phase is completely independent of Switchyard. Any framework that produces a loadable checkpoint and can expose it via an HTTP service satisfies the requirements.

## Step 2: Deploy the Fine-Tuned Model Behind an API Endpoint

After training, serve the checkpoint using an LLM inference engine that supports OpenAI-compatible chat completion formats. Suitable options include **vLLM**, **Text Generation Inference (TGI)**, or custom FastAPI wrappers.

Ensure your deployment exposes a `base_url` and accepts authentication via environment variables (e.g., `MY_SERVER_API_KEY`), as Switchyard's TOML configuration references these credentials when initializing LLM clients.

## Step 3: Configure the Fine-Tuned Model in routes.toml

Switchyard centralizes routing logic in a TOML configuration file. Define an LLM client pointing to your deployed endpoint, then register it as a target for routing algorithms:

```toml
schema_version = 1

# LLM client definition (replace with your server’s URL)

[llm_clients.my_server]
format = "openai_chat"
base_url = "https://my-finetuned-model.example.com/v1"
api_key_env = "MY_SERVER_API_KEY"

# Target that uses the fine‑tuned model

[targets.fine_tuned]
id = "myorg/fine‑tuned‑model"
llm_client = "my_server"

# Example routing algorithm that prefers a cheap model but falls back to the fine‑tuned one

[routes.switchyard]
type = "stage_router"
capable_target = "fine_tuned"      # the fine‑tuned model

efficient_target = "efficient"    # a cheaper fallback model (defined elsewhere)

picker = "efficient_first"
confidence_threshold = 0.5

```

The `schema_version = 1` declaration enables Switchyard's configuration parser—validated in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs)—to process routes, clients, and targets. For complete schema specifications, reference [`docs/reference/toml_schema.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/reference/toml_schema.md) in the NVIDIA-NeMo/Switchyard repository.

## Step 4: Drive Routing Logic with the libsy Python Library

Switchyard exposes its routing algorithms through the `libsy` Python package. The following implementation uses the **stage_router** algorithm—defined in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs)—to route complex requests to your fine-tuned model while falling back to efficient baselines for simpler queries:

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build the routing algorithm

algorithm = stage_router(
    capable="fine_tuned",
    efficient="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Call the underlying model client (e.g., an OpenAI‑compatible client)

                resp = await clients[call.models[0]].call({**call.request, "model": call.models[0]})
                call.respond(LlmResponse.Agg(resp))
            case Step.Done(outcome):
                return outcome.response or await clients[outcome.selected_model_ids[0]].call(outcome.request)
    raise RuntimeError("Algorithm finished without a decision")

```

This pattern isolates your application from specific model implementations. When deploying an updated fine-tuned version, modify [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) to reference the new endpoint without changing Python routing code. See [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) in the repository for a complete runnable implementation.

## Summary

- **Switchyard does not train models**: It is a pure routing layer requiring externally fine-tuned checkpoints served via HTTP endpoints.
- **Configuration-driven integration**: Register fine-tuned models as targets in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) by defining LLM clients that connect to OpenAI-compatible endpoints.
- **Algorithmic selection**: Use built-in algorithms like `stage_router` (implemented in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs)) to intelligently balance between fine-tuned capable models and cost-efficient baselines.
- **Zero application code changes**: Swapping fine-tuned models requires only TOML updates, not modifications to your Python or Rust application logic.
- **Flexible deployment**: Any serving framework producing OpenAI-compatible APIs—including vLLM, TGI, or custom servers—integrates seamlessly with Switchyard's client abstraction.

## Frequently Asked Questions

### Can Switchyard fine-tune models directly?

No. Switchyard is strictly an LLM routing library. According to the source code in [`crates/switchyard-runner/src/runner.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/runner.rs), the system handles request distribution and algorithm execution, not gradient computation or weight updates. You must use external frameworks like NVIDIA NeMo, Hugging Face Transformers, or PyTorch for all fine-tuning operations.

### What API formats must my fine-tuned model expose?

Switchyard requires OpenAI-compatible chat completion endpoints or Anthropic-style API shapes. Most modern deployment frameworks—including vLLM, Text Generation Inference, and OpenRouter—provide this compatibility. The `format` key in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml) (e.g., `format = "openai_chat"`) instructs Switchyard how to serialize requests to your specific endpoint.

### How do I route specific requests to my fine-tuned model versus a base model?

Configure a routing algorithm such as `stage_router` in [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml). Set `capable_target` to your fine-tuned model's target ID and `efficient_target` to a base model. The algorithm—implemented in [`crates/libsy/src/algorithms/stage.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/stage.rs)—evaluates confidence thresholds and query complexity to determine whether to invoke the fine-tuned model or the efficient fallback.

### Do I need to restart Switchyard when updating fine-tuned models?

No. Switchyard loads routing configurations dynamically from [`routes.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routes.toml). To deploy a new fine-tuned version, update the `base_url` in the `[llm_clients]` section or create a new target, then reference it in your routing rules. The runner process detects these changes without requiring application restarts or code modifications.