How to Install NVIDIA Switchyard: Python Package, Rust Server, and Library Integration

Install NVIDIA Switchyard via pip install nemo-switchyard for Python bindings, cargo install --locked switchyard-server for the standalone HTTP server, or add individual crates to your Cargo.toml for embedded Rust applications.

NVIDIA Switchyard is a hybrid Python‑Rust orchestration layer for large‑language‑model (LLM) traffic. According to the NVIDIA‑NeMo/Switchyard source code, the project separates concerns into three installable components: Python bindings that expose routing algorithms via PyO3, a production‑grade Rust server binary, and modular Rust crates for custom integrations. Depending on your use case—whether you need a Python library, a dedicated proxy server, or embedded routing logic—you will follow different installation paths.

Prerequisites

Before installing any component, verify your environment meets the minimum requirements. The Python package requires Python ≥ 3.10, while the Rust server and libraries require Rust ≥ 1.96.1. Hardware constraints apply to pre‑built wheels: Linux x86_64 wheels require a CPU with AVX2‑class instructions, and aarch64 wheels require a Neoverse N1‑class CPU.

Install the Python Package

The simplest way to use Switchyard in Python applications is via the nemo-switchyard package published on PyPI. This package bundles the native Rust extension and has no runtime dependencies beyond the compiled binary.

Run the following command:

pip install nemo-switchyard

After installation, import the top‑level switchyard module to access the libsy routing algorithms. Because the native extension is bundled, you do not need a separate Rust toolchain for basic Python usage. The entry point is defined in pyproject.toml, with core algorithm implementations located in switchyard/libsy/algorithms.py.

Install the Standalone Rust Server

For production deployments requiring a dedicated LLM routing service, install the switchyard-server binary from crates.io. This server loads TOML deployment configurations and proxies requests to upstream LLM providers.

Execute:

cargo install --locked switchyard-server

Validate the installation with a dry‑run:

switchyard-server --config routes.toml --dry-run

Start the server on your desired port (for example, 4000):

switchyard-server --config routes.toml --port 4000

Configuration details and deployment schemas are documented in crates/switchyard-server/README.md and the Getting Started guide at docs/getting_started.md.

Use the Rust Libraries in Custom Applications

To embed Switchyard’s routing logic directly into your own Rust binaries, add the required crates to your Cargo.toml:

[dependencies]
switchyard-libsy = "0.2.0"
switchyard-protocol = "0.2.0"
switchyard-llm-client = "0.2.0"
switchyard-translation = "0.2.0"

These crates provide distinct functionality: switchyard-libsy contains the core routing algorithms (such as RoundRobin), switchyard-protocol defines provider‑neutral request and response types, switchyard-llm-client handles HTTP translation logic, and switchyard-translation manages wire‑format codecs. Source files for these components reside in the respective crates/ directory of the repository.

Development Setup (Optional)

When contributing to Switchyard or building from source, clone the repository and run the following commands. The dev dependency group contains linting, formatting, and test utilities that are excluded from the published wheel.

uv sync                     # Install Python dev dependencies

uv run maturin develop      # Build the native PyO3 extension

cargo test --workspace      # Run Rust unit tests

uv run pytest tests/ -v     # Run Python test suite

Verification and Usage Examples

After installation, verify functionality using the patterns below.

Python API Example

The following snippet demonstrates routing a request through the Python bindings. The switchyard.libsy.Algorithms class exposes native Rust routing logic compiled via PyO3.

import switchyard

# Initialize a round-robin router

router = switchyard.libsy.Algorithms.round_robin()

# Construct a provider-neutral request

request = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello, Switchyard!"}]
}

# Route the request; the library handles HTTP translation internally

response = router.route(request)
print(response["choices"][0]["message"]["content"])

Server and cURL Example

Start the standalone server and query it with an OpenAI‑compatible payload. This workflow is documented in docs/getting_started.md under the Server Path section.


# Start the server in the background

switchyard-server --config routes.toml --port 4000 &

# Send a chat completion request

curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role":"user","content":"What is Switchyard?"}]
      }'

Rust Library Example

Embed the routing algorithm directly in a Rust application. The source for RoundRobin is located in crates/switchyard-libsy/src/algorithms.rs.

use switchyard_libsy::algorithms::RoundRobin;
use switchyard_protocol::ChatRequest;

fn main() {
    let router = RoundRobin::new();
    let request = ChatRequest {
        model: "gpt-4o-mini".into(),
        messages: vec![("user".into(), "Hello from Rust!".into())],
        ..Default::default()
    };
    let response = router.route(request);
    println!("{:?}", response);
}

Summary

  • Python users should run pip install nemo-switchyard to receive the bundled PyO3 extension with no additional Rust toolchain required.
  • Server operators should run cargo install --locked switchyard-server to deploy the standalone HTTP proxy defined in crates/switchyard-server/.
  • Rust developers should add specific crates—switchyard-libsy, switchyard-protocol, switchyard-llm-client, and switchyard-translation—to Cargo.toml for embedded use.
  • Contributors must install Python ≥ 3.10 and Rust ≥ 1.96.1, then use uv and cargo to build and test the hybrid codebase.

Frequently Asked Questions

Do I need to install Rust to use the Python package?

No. The nemo-switchyard wheel distributed on PyPI contains the pre‑compiled Rust extension bundled via PyO3. You only need Python ≥ 3.10 and a compatible CPU (AVX2 for x86_64, Neoverse N1 for aarch64). A Rust toolchain is only required if you intend to build the package from source.

What is the difference between the Python bindings and the standalone server?

The Python bindings expose Switchyard’s routing algorithms as an importable library for embedding in Python applications, as implemented in switchyard/libsy/algorithms.py. The standalone server (switchyard-server) is a production‑grade Rust binary that runs as a separate process, loads TOML configurations, and exposes an HTTP interface for proxying LLM requests. Choose the former for library integration and the latter for infrastructure‑level traffic management.

Can I install the server from PyPI instead of crates.io?

No. The nemo-switchyard PyPI package contains the Python bindings and native extension, but it does not install the switchyard-server binary. To run the standalone server, you must use cargo install --locked switchyard-server or build the binary from the crates/switchyard-server/ source directory.

Which crate provides the core routing algorithms?

The switchyard-libsy crate (version 0.2.0) contains the core routing logic, including implementations like RoundRobin. This crate is used internally by the Python bindings via PyO3 and can be linked directly into custom Rust applications by adding it to Cargo.toml.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →