# Switchyard Library APIs: Server Management and LLM Routing Reference

> Explore NVIDIA Switchyard library APIs for server management and LLM routing. Discover key Python packages switchyard_rust.server and switchyard.libsy for building routing algorithms.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: api-reference
- Published: 2026-08-23

---

**The Switchyard library exposes its core functionality through two primary Python packages: `switchyard_rust.server` for managing the native Rust HTTP server lifecycle, and `switchyard_rust.libsy` (conveniently re-exported as `switchyard.libsy`) for constructing routing algorithms including LLM classifiers and stage routers.**

Switchyard is a dual-language library developed by NVIDIA that lets Python applications orchestrate Large Language Model traffic while delegating high-performance routing and serving logic to a native Rust implementation. Mastering the Switchyard library APIs enables developers to programmatically control server initialization, manage model endpoints, and implement sophisticated routing strategies for intelligent traffic distribution.

## Server API: Managing the Native Rust Runtime

The `switchyard_rust.server` module provides the primary interface for starting and managing the native Switchyard server. The `Server` class acts as a thin Python wrapper around the compiled Rust binary, loading the native library on demand and providing a Pythonic context manager for lifecycle management.

### Starting the Server with Configuration

Instantiate the `Server` class by passing a TOML deployment configuration (or file path) and optional port parameter. Setting `port=0` instructs the operating system to select an available ephemeral port automatically.

```python
from switchyard_rust.server import Server

# Load TOML configuration and start the server

with Server("routes.toml", port=0) as server:
    print(f"Running on http://{server.base_url}")
    # Client code can now call OpenAI-compatible endpoints

```

According to the source code in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py), the `Server` class handles native library loading and exposes the OpenAI/Anthropic-compatible HTTP endpoints through the underlying Rust binary.

### Server Instance Methods and Properties

The `Server` instance provides several key attributes and methods for runtime inspection and control:

- **`port`** – Returns the actual port number the server is listening on (particularly useful when `port=0` was specified).
- **`base_url`** – Provides a convenient base URL string for constructing client requests.
- **`caller_auth_kind(model)`** – Returns the required authentication kind for a specific model name.
- **`close(timeout_secs)`** – Initiates a graceful shutdown of the server with the specified timeout.

## Algorithm API: Implementing Routing Logic

The `switchyard_rust.libsy` module contains the core routing functionality executed by the Rust `libsy` crate. All routing algorithms implement the `Algorithm` interface, which defines the contract for processing LLM requests and determining target models.

### The Algorithm Class and run_stream Method

Every algorithm instance implements the `run_stream` method, which returns an async iterator yielding routing decisions. The method signature accepts request metadata and optional headers:

```python
algorithm.run_stream(
    request: Mapping[str, object],
    headers: Mapping[str, str] | None = None
) -> AsyncIterator[Step.CallModel | Step.Done]

```

As implemented in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), the iterator yields:
- **`Step.CallModel`** – Indicates which model(s) to invoke for the current request.
- **`Step.Done`** – Signifies the final routing outcome with the selected model ID.

## Routing Algorithms and Classifiers

The Switchyard library APIs provide several factory functions for constructing routing algorithms tailored to different operational requirements.

### Capability Classifier

The capability classifier selects a target based on task capability predictions, routing to an efficient model for simple tasks and a capable model for complex ones. Configure it using `LlmClassifierConfig.capability()`:

```python
from switchyard.libsy import llm_classifier, LlmClassifierConfig, TaskClassifierConfig

cfg = LlmClassifierConfig.capability(
    judge_target="classifier",
    efficient_target="gpt-3.5",
    capable_target="gpt-4",
    config=TaskClassifierConfig(
        base_threshold=0.5,
        session_affinity=True,
    ),
)
alg = llm_classifier(cfg)

```

### Escalation Classifier

The escalation classifier first attempts to use an efficient model, then escalates to a capable model if the response suggests insufficient performance. Configuration uses `LlmClassifierConfig.escalation()`:

```python
from switchyard.libsy import llm_classifier, LlmClassifierConfig, EscalationClassifierConfig

cfg = LlmClassifierConfig.escalation(
    judge_target="classifier",
    efficient_target="gpt-3.5",
    capable_target="gpt-4",
    config=EscalationClassifierConfig(confirmations=2, recent_turn_window=28),
)
alg = llm_classifier(cfg)

```

### Custom Classifier

For domain-specific routing, the custom classifier routes among user-defined targets using JSON-schema-validated labels:

```python
from switchyard.libsy import llm_classifier, LlmClassifierConfig, CustomClassifierConfig

cfg = LlmClassifierConfig.custom(
    judge_target="classifier",
    targets=[("fast", "gpt-3.5"), ("smart", "gpt-4")],
    default_target="fast",
    config=CustomClassifierConfig(
        prompt="Classify the task",
        response_schema={"type": "string", "enum": ["fast", "smart"]},
        selector="label",
    ),
)
alg = llm_classifier(cfg)

```

All classifier configurations are defined in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), which provides the binding layer to the Rust implementation.

### Stage Router API

The stage router algorithm determines **when** to switch from an efficient target to a capable one based on classifier confidence scores. The `stage_router` factory function creates this algorithm:

```python
from switchyard.libsy import stage_router

alg = stage_router(
    capable_target="gpt-4",
    efficient_target="gpt-3.5",
    picker="max_confidence",
    confidence_threshold=0.8,
    recent_window=10,
    escalation_note="Escalated due to low confidence",
)

```

### Utility Algorithms

For testing and experimental workloads, [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) provides additional helpers:

- **`noop()`** – Returns a no-op algorithm that immediately returns a deterministic outcome without invoking any models.
- **`random(targets, weights=None, seed=None)`** – Selects targets uniformly or with specified weights for A/B testing scenarios.

## Import Convenience Layer

To simplify client code, the `switchyard.libsy.algorithms` module re-exports all factory functions from the underlying Rust bindings. This allows users to import directly from the top-level `switchyard` package rather than referencing the internal `switchyard_rust` structure:

```python
from switchyard.libsy import llm_classifier, stage_router, LlmClassifierConfig

```

The re-export definitions reside in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), which serves as the canonical import location for Python applications.

## End-to-End Integration Example

The following example demonstrates the complete workflow: starting a server, configuring a capability classifier, and processing a request through the algorithm:

```python
import asyncio
from switchyard_rust.server import Server
from switchyard.libsy import llm_classifier, LlmClassifierConfig, TaskClassifierConfig

async def main():
    # Initialize server with TOML configuration

    server = Server("examples/routes.toml", port=0)
    await server.__aenter__()
    
    try:
        # Configure capability-based routing

        clf_cfg = LlmClassifierConfig.capability(
            judge_target="classifier",
            efficient_target="gpt-3.5",
            capable_target="gpt-4",
            config=TaskClassifierConfig(base_threshold=0.6, session_affinity=True),
        )
        algorithm = llm_classifier(clf_cfg)
        
        # Execute routing decision

        request = {
            "model": "gpt-3.5",
            "messages": [{"role": "user", "content": "Summarize Shakespeare"}]
        }
        
        async for step in algorithm.run_stream(request):
            if isinstance(step, Step.CallModel):
                print("Calling model:", step.call.models)
            else:  # Step.Done

                print("Final selection:", step.outcome.selected_model_id)
    finally:
        await server.__aexit__(None, None, None)

asyncio.run(main())

```

In this implementation, the `Server` object manages the model lifecycle and HTTP endpoints, while the `Algorithm` operates independently to determine routing decisions through the `run_stream` async iterator.

## Summary

- **Server API** ([`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py)): Use the `Server` class to start the native Rust HTTP server, manage ports, and handle authentication configurations through a Pythonic context manager interface.
- **Algorithm API** ([`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py)): Implement routing logic using factory functions like `llm_classifier()` and `stage_router()`, which return `Algorithm` instances that yield routing decisions via `run_stream()`.
- **Classifier Types**: Choose between capability classifiers (task-based routing), escalation classifiers (retry-based routing), and custom classifiers (schema-validated routing) based on your traffic distribution requirements.
- **Convenience Imports**: Import from `switchyard.libsy` rather than `switchyard_rust.libsy` for cleaner dependency management in production applications.

## Frequently Asked Questions

### What is the difference between switchyard_rust.server and switchyard.libsy?

**`switchyard_rust.server`** manages the HTTP server lifecycle and OpenAPI-compatible endpoints, wrapping the native Rust binary in a Python class. **`switchyard.libsy`** (re-exported from `switchyard_rust.libsy`) provides the routing algorithm interface for determining which LLM should handle specific requests. The server manages *how* requests are received, while the algorithms determine *where* requests are sent.

### How do I configure a custom routing algorithm in Switchyard?

Import `LlmClassifierConfig.custom()` from `switchyard.libsy` to define targets with custom labels and JSON schemas. Specify your targets as tuples of `(label, model_name)`, provide a validation schema, and set a default target for fallbacks. The classifier will validate responses against your schema and route accordingly.

### Can I run the Switchyard server without using the context manager?

Yes, though the context manager (`with Server(...) as server:`) is recommended for automatic cleanup. You can manually instantiate `Server("config.toml", port=0)`, call `await server.__aenter__()` to start, and explicitly call `await server.__aexit__(None, None, None)` or `server.close(timeout_secs)` when shutting down.

### What data does the Algorithm.run_stream method return?

The `run_stream` method returns an async iterator that yields `Step` objects from the Rust runtime. Each iteration yields either a `Step.CallModel` containing the target model identifier and call parameters, or a `Step.Done` object containing the final routing outcome including the selected model ID. Your application logic must handle these steps to execute the model calls and return responses.