How to Fine-Tune Models with Switchyard: A Complete Integration Guide

To fine-tune models with Switchyard, train externally using NVIDIA NeMo or similar frameworks, deploy the checkpoint behind an OpenAI-compatible endpoint, register it as a target in routes.toml, and route traffic using Switchyard's algorithms.

Switchyard is an LLM routing library—not a training framework—that intelligently distributes requests across multiple language models. According to the NVIDIA-NeMo/Switchyard source code, you complete fine-tuning workflows outside the system, then integrate the deployed models into Switchyard's configuration-driven routing layer. This architecture decouples model development from production traffic management, allowing you to swap fine-tuned checkpoints without modifying application code.

Understanding Switchyard's Architecture

Switchyard Is a Router, Not a Trainer

Switchyard functions exclusively as a request routing engine. As implemented in crates/switchyard-runner/src/runner.rs, the runtime parses TOML configurations and drives algorithms like stage_router, but performs no backpropagation or weight updates. The system expects all models—including fine-tuned variants—to be served via external HTTP endpoints that expose LLM-compatible APIs.

Step 1: Fine-Tune Your Model Using NVIDIA NeMo

Since Switchyard does not provide training capabilities, use NVIDIA NeMo or any ML framework capable of exporting to a servable format. The following example demonstrates supervised fine-tuning using NeMo's prompt learning model:


# Example using NVIDIA‑NeMo (requires nemo‑core & nemo‑model‑zoo)

from nemo.collections.nlp.models.language_modeling import GPTPromptLearningModel

model = GPTPromptLearningModel.restore_from(
    "base_model.nemo"
).to("cuda")

# Prepare your SFT dataset (json / jsonl with {"prompt":..., "completion":...})

model.train_dataset = "sft_dataset.jsonl"

trainer = nemo.utils.Trainer(
    max_epochs=3,
    gpus=1,
    precision=16,
)
trainer.fit(model)

# Export the fine‑tuned checkpoint

model.save_to("fine_tuned_model.nemo")

This training phase is completely independent of Switchyard. Any framework that produces a loadable checkpoint and can expose it via an HTTP service satisfies the requirements.

Step 2: Deploy the Fine-Tuned Model Behind an API Endpoint

After training, serve the checkpoint using an LLM inference engine that supports OpenAI-compatible chat completion formats. Suitable options include vLLM, Text Generation Inference (TGI), or custom FastAPI wrappers.

Ensure your deployment exposes a base_url and accepts authentication via environment variables (e.g., MY_SERVER_API_KEY), as Switchyard's TOML configuration references these credentials when initializing LLM clients.

Step 3: Configure the Fine-Tuned Model in routes.toml

Switchyard centralizes routing logic in a TOML configuration file. Define an LLM client pointing to your deployed endpoint, then register it as a target for routing algorithms:

schema_version = 1

# LLM client definition (replace with your server’s URL)

[llm_clients.my_server]
format = "openai_chat"
base_url = "https://my-finetuned-model.example.com/v1"
api_key_env = "MY_SERVER_API_KEY"

# Target that uses the fine‑tuned model

[targets.fine_tuned]
id = "myorg/fine‑tuned‑model"
llm_client = "my_server"

# Example routing algorithm that prefers a cheap model but falls back to the fine‑tuned one

[routes.switchyard]
type = "stage_router"
capable_target = "fine_tuned"      # the fine‑tuned model

efficient_target = "efficient"    # a cheaper fallback model (defined elsewhere)

picker = "efficient_first"
confidence_threshold = 0.5

The schema_version = 1 declaration enables Switchyard's configuration parser—validated in crates/switchyard-runner/src/runner.rs—to process routes, clients, and targets. For complete schema specifications, reference docs/reference/toml_schema.md in the NVIDIA-NeMo/Switchyard repository.

Step 4: Drive Routing Logic with the libsy Python Library

Switchyard exposes its routing algorithms through the libsy Python package. The following implementation uses the stage_router algorithm—defined in crates/libsy/src/algorithms/stage.rs—to route complex requests to your fine-tuned model while falling back to efficient baselines for simpler queries:

from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build the routing algorithm

algorithm = stage_router(
    capable="fine_tuned",
    efficient="efficient",
    picker="efficient_first",
    confidence_threshold=0.5,
)

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    async for step in algorithm.run_stream(request):
        match step:
            case Step.CallModel(call):
                # Call the underlying model client (e.g., an OpenAI‑compatible client)

                resp = await clients[call.models[0]].call({**call.request, "model": call.models[0]})
                call.respond(LlmResponse.Agg(resp))
            case Step.Done(outcome):
                return outcome.response or await clients[outcome.selected_model_ids[0]].call(outcome.request)
    raise RuntimeError("Algorithm finished without a decision")

This pattern isolates your application from specific model implementations. When deploying an updated fine-tuned version, modify routes.toml to reference the new endpoint without changing Python routing code. See examples/libsy.py in the repository for a complete runnable implementation.

Summary

  • Switchyard does not train models: It is a pure routing layer requiring externally fine-tuned checkpoints served via HTTP endpoints.
  • Configuration-driven integration: Register fine-tuned models as targets in routes.toml by defining LLM clients that connect to OpenAI-compatible endpoints.
  • Algorithmic selection: Use built-in algorithms like stage_router (implemented in crates/libsy/src/algorithms/stage.rs) to intelligently balance between fine-tuned capable models and cost-efficient baselines.
  • Zero application code changes: Swapping fine-tuned models requires only TOML updates, not modifications to your Python or Rust application logic.
  • Flexible deployment: Any serving framework producing OpenAI-compatible APIs—including vLLM, TGI, or custom servers—integrates seamlessly with Switchyard's client abstraction.

Frequently Asked Questions

Can Switchyard fine-tune models directly?

No. Switchyard is strictly an LLM routing library. According to the source code in crates/switchyard-runner/src/runner.rs, the system handles request distribution and algorithm execution, not gradient computation or weight updates. You must use external frameworks like NVIDIA NeMo, Hugging Face Transformers, or PyTorch for all fine-tuning operations.

What API formats must my fine-tuned model expose?

Switchyard requires OpenAI-compatible chat completion endpoints or Anthropic-style API shapes. Most modern deployment frameworks—including vLLM, Text Generation Inference, and OpenRouter—provide this compatibility. The format key in routes.toml (e.g., format = "openai_chat") instructs Switchyard how to serialize requests to your specific endpoint.

How do I route specific requests to my fine-tuned model versus a base model?

Configure a routing algorithm such as stage_router in routes.toml. Set capable_target to your fine-tuned model's target ID and efficient_target to a base model. The algorithm—implemented in crates/libsy/src/algorithms/stage.rs—evaluates confidence thresholds and query complexity to determine whether to invoke the fine-tuned model or the efficient fallback.

Do I need to restart Switchyard when updating fine-tuned models?

No. Switchyard loads routing configurations dynamically from routes.toml. To deploy a new fine-tuned version, update the base_url in the [llm_clients] section or create a new target, then reference it in your routing rules. The runner process detects these changes without requiring application restarts or code modifications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →