# Are There Pre-Trained Models Available for Switchyard? A Technical Guide to LLM Routing

> Discover if Switchyard offers pre-trained models. Learn how this LLM routing layer connects to external services like OpenAI and NVIDIA NeMo Relay for efficient inference.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-11

---

**Switchyard does not ship pre-trained models; it is a lightweight routing layer that forwards inference requests to externally hosted LLM services such as OpenAI, Anthropic, and NVIDIA NeMo Relay.**

Switchyard is an open-source inference router from the NVIDIA-NeMo organization designed to optimize traffic across multiple large language model (LLM) providers. Unlike model repositories that distribute `.ckpt` or `.bin` files, Switchyard contains only Rust source code, Python bindings, and configuration utilities that dynamically route prompts to pre-trained models hosted by third parties.

## Switchyard is a Router, Not a Model Repository

The Switchyard repository contains **no model assets**. According to the source tree analysis, the codebase consists exclusively of routing logic, HTTP clients, and provider plugins. There are no transformer weights, checkpoints, or embedding matrices bundled in the distribution.

In [`crates/libsy-llm-client/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md), the project explicitly defines itself as an HTTP client bridge to external LLM APIs. The library implements the `libsy` routing algorithm to select optimal backend models at runtime, but it does not embed any generative AI capabilities itself. When you deploy Switchyard, you must configure it to point at existing model endpoints that already host pre-trained weights.

## Configuring External Pre-Trained Models via TOML

Switchyard uses declarative TOML configuration files to register available models from external providers. The router references this configuration to determine which backend serves a given request based on latency, cost, or availability preferences.

A typical configuration from [`examples/litellm/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/README.md) demonstrates how to register GPT-4o and Claude models without local storage:

```toml

# examples/litellm/config.toml

[providers.openai]
api_key = "YOUR_OPENAI_API_KEY"
models = ["gpt-4o", "gpt-3.5-turbo"]

[providers.anthropic]
api_key = "YOUR_ANTHROPIC_API_KEY"
models = ["claude-3-sonnet-20240229"]

[routing]

# Prefer lower latency, then cost

preferences = ["latency", "cost"]

```

This configuration instructs Switchyard to route inference requests to OpenAI's and Anthropic's pre-trained models through their respective APIs. The `models` array contains string identifiers for cloud-hosted checkpoints, not local file paths.

## Integrating with LLM Providers: NeMo Relay and LiteLLM

Switchyard supports provider-specific plugins that wrap external serving infrastructure. These plugins translate Switchyard's internal request format into provider-specific API calls, allowing the router to leverage pre-trained models from diverse sources without native code changes.

### NVIDIA NeMo Relay Plugin

The [`crates/switchyard-nemo-relay-plugin/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/README.md) demonstrates routing to NVIDIA-hosted models via a local or remote NeMo endpoint:

```python
from switchyard_nemo_relay import NemoRelayPlugin
from switchyard import Runner

plugin = NemoRelayPlugin(
    endpoint="http://localhost:8000",
    model_name="nvidia/megatron-gpt-2b",
)
runner = Runner(config, plugins=[plugin])

```

Even when targeting `localhost`, the pre-trained weights reside in the NeMo Relay service, not within Switchyard's process space. Switchyard merely constructs the HTTP request and forwards the prompt to the inference server.

### LiteLLM Integration

The LiteLLM plugin (referenced in [`examples/litellm/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/litellm/README.md)) provides a unified interface to multiple commercial APIs. Switchyard treats LiteLLM as a backend provider, routing requests through it to access models like GPT-4o or Claude 3 without direct provider integration.

## Implementation: Routing Requests to Remote Models

Once configured, Switchyard's Python bindings load the TOML specification and instantiate a `Runner` object to manage request dispatch. The `run_stream()` method sends prompts to the selected external model based on the routing algorithm defined in `libsy`.

```python

# switchyard_rust/libsy.py

from switchyard import Runner, Config

# Load the TOML configuration containing external model references

cfg = Config.from_toml("examples/litellm/config.toml")
runner = Runner(cfg)

# Route to the best available pre-trained model

response = runner.run_stream(
    prompt="Explain quantum entanglement in simple terms.",
    max_tokens=200,
)
print(response.text)

```

In this execution flow, `Config.from_toml()` parses the provider credentials and model identifiers, while `runner.run_stream()` selects the optimal backend (e.g., OpenAI's GPT-4o or Anthropic's Claude) and forwards the request. The response text originates from the external provider's pre-trained model, streamed back through Switchyard's routing layer.

## Summary

- **Switchyard contains no model weights**: The repository includes only Rust code for routing logic and Python bindings for configuration; there are no `.ckpt`, `.bin`, or Safetensors files.
- **Provider-driven architecture**: Pre-trained models must be hosted externally by OpenAI, Anthropic, NVIDIA NeMo Relay, or compatible API endpoints.
- **TOML-based configuration**: You define available models and routing preferences in configuration files, not by downloading weights.
- **Plugin extensibility**: Integration plugins like `switchyard-nemo-relay` connect to pre-trained models served via HTTP APIs, whether cloud-hosted or local inference servers.

## Frequently Asked Questions

### Does Switchyard include any pre-trained model files in the repository?

No. Switchyard is strictly an inference router. The source tree contains only Rust crates for routing logic, Python bindings, and configuration examples. All model assets—whether GPT-4o, Claude 3, or Megatron checkpoints—reside with external providers.

### How do I use GPT-4 or Claude models with Switchyard?

Configure the `providers.openai` or `providers.anthropic` sections in your TOML configuration file with your API keys and desired model identifiers (e.g., `gpt-4o`, `claude-3-sonnet-20240229`). Switchyard will route eligible requests to these endpoints using the `libsy` routing algorithm.

### Can I use Switchyard with locally hosted pre-trained models?

Yes, if the local models are served via an HTTP API endpoint. The NeMo Relay plugin, for example, connects to `localhost:8000` to access locally served NVIDIA models. Switchyard routes requests to this endpoint, but the model weights remain loaded in the external inference server, not inside Switchyard itself.

### What is the `libsy-llm-client` crate used for?

The `libsy-llm-client` crate (documented in [`crates/libsy-llm-client/README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy-llm-client/README.md)) implements the HTTP client layer that communicates with external LLM APIs. It handles authentication, request serialization, and response streaming from pre-trained model providers to Switchyard's routing core.