How to Install NVIDIA Switchyard: Python Package, Rust Server, and Library Integration
Install NVIDIA Switchyard via pip install nemo-switchyard for Python bindings, cargo install --locked switchyard-server for the standalone HTTP server, or add individual crates to your Cargo.toml for embedded Rust applications.
NVIDIA Switchyard is a hybrid Python‑Rust orchestration layer for large‑language‑model (LLM) traffic. According to the NVIDIA‑NeMo/Switchyard source code, the project separates concerns into three installable components: Python bindings that expose routing algorithms via PyO3, a production‑grade Rust server binary, and modular Rust crates for custom integrations. Depending on your use case—whether you need a Python library, a dedicated proxy server, or embedded routing logic—you will follow different installation paths.
Prerequisites
Before installing any component, verify your environment meets the minimum requirements. The Python package requires Python ≥ 3.10, while the Rust server and libraries require Rust ≥ 1.96.1. Hardware constraints apply to pre‑built wheels: Linux x86_64 wheels require a CPU with AVX2‑class instructions, and aarch64 wheels require a Neoverse N1‑class CPU.
Install the Python Package
The simplest way to use Switchyard in Python applications is via the nemo-switchyard package published on PyPI. This package bundles the native Rust extension and has no runtime dependencies beyond the compiled binary.
Run the following command:
pip install nemo-switchyard
After installation, import the top‑level switchyard module to access the libsy routing algorithms. Because the native extension is bundled, you do not need a separate Rust toolchain for basic Python usage. The entry point is defined in pyproject.toml, with core algorithm implementations located in switchyard/libsy/algorithms.py.
Install the Standalone Rust Server
For production deployments requiring a dedicated LLM routing service, install the switchyard-server binary from crates.io. This server loads TOML deployment configurations and proxies requests to upstream LLM providers.
Execute:
cargo install --locked switchyard-server
Validate the installation with a dry‑run:
switchyard-server --config routes.toml --dry-run
Start the server on your desired port (for example, 4000):
switchyard-server --config routes.toml --port 4000
Configuration details and deployment schemas are documented in crates/switchyard-server/README.md and the Getting Started guide at docs/getting_started.md.
Use the Rust Libraries in Custom Applications
To embed Switchyard’s routing logic directly into your own Rust binaries, add the required crates to your Cargo.toml:
[dependencies]
switchyard-libsy = "0.2.0"
switchyard-protocol = "0.2.0"
switchyard-llm-client = "0.2.0"
switchyard-translation = "0.2.0"
These crates provide distinct functionality: switchyard-libsy contains the core routing algorithms (such as RoundRobin), switchyard-protocol defines provider‑neutral request and response types, switchyard-llm-client handles HTTP translation logic, and switchyard-translation manages wire‑format codecs. Source files for these components reside in the respective crates/ directory of the repository.
Development Setup (Optional)
When contributing to Switchyard or building from source, clone the repository and run the following commands. The dev dependency group contains linting, formatting, and test utilities that are excluded from the published wheel.
uv sync # Install Python dev dependencies
uv run maturin develop # Build the native PyO3 extension
cargo test --workspace # Run Rust unit tests
uv run pytest tests/ -v # Run Python test suite
Verification and Usage Examples
After installation, verify functionality using the patterns below.
Python API Example
The following snippet demonstrates routing a request through the Python bindings. The switchyard.libsy.Algorithms class exposes native Rust routing logic compiled via PyO3.
import switchyard
# Initialize a round-robin router
router = switchyard.libsy.Algorithms.round_robin()
# Construct a provider-neutral request
request = {
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello, Switchyard!"}]
}
# Route the request; the library handles HTTP translation internally
response = router.route(request)
print(response["choices"][0]["message"]["content"])
Server and cURL Example
Start the standalone server and query it with an OpenAI‑compatible payload. This workflow is documented in docs/getting_started.md under the Server Path section.
# Start the server in the background
switchyard-server --config routes.toml --port 4000 &
# Send a chat completion request
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role":"user","content":"What is Switchyard?"}]
}'
Rust Library Example
Embed the routing algorithm directly in a Rust application. The source for RoundRobin is located in crates/switchyard-libsy/src/algorithms.rs.
use switchyard_libsy::algorithms::RoundRobin;
use switchyard_protocol::ChatRequest;
fn main() {
let router = RoundRobin::new();
let request = ChatRequest {
model: "gpt-4o-mini".into(),
messages: vec![("user".into(), "Hello from Rust!".into())],
..Default::default()
};
let response = router.route(request);
println!("{:?}", response);
}
Summary
- Python users should run
pip install nemo-switchyardto receive the bundled PyO3 extension with no additional Rust toolchain required. - Server operators should run
cargo install --locked switchyard-serverto deploy the standalone HTTP proxy defined incrates/switchyard-server/. - Rust developers should add specific crates—
switchyard-libsy,switchyard-protocol,switchyard-llm-client, andswitchyard-translation—toCargo.tomlfor embedded use. - Contributors must install Python ≥ 3.10 and Rust ≥ 1.96.1, then use
uvandcargoto build and test the hybrid codebase.
Frequently Asked Questions
Do I need to install Rust to use the Python package?
No. The nemo-switchyard wheel distributed on PyPI contains the pre‑compiled Rust extension bundled via PyO3. You only need Python ≥ 3.10 and a compatible CPU (AVX2 for x86_64, Neoverse N1 for aarch64). A Rust toolchain is only required if you intend to build the package from source.
What is the difference between the Python bindings and the standalone server?
The Python bindings expose Switchyard’s routing algorithms as an importable library for embedding in Python applications, as implemented in switchyard/libsy/algorithms.py. The standalone server (switchyard-server) is a production‑grade Rust binary that runs as a separate process, loads TOML configurations, and exposes an HTTP interface for proxying LLM requests. Choose the former for library integration and the latter for infrastructure‑level traffic management.
Can I install the server from PyPI instead of crates.io?
No. The nemo-switchyard PyPI package contains the Python bindings and native extension, but it does not install the switchyard-server binary. To run the standalone server, you must use cargo install --locked switchyard-server or build the binary from the crates/switchyard-server/ source directory.
Which crate provides the core routing algorithms?
The switchyard-libsy crate (version 0.2.0) contains the core routing logic, including implementations like RoundRobin. This crate is used internally by the Python bindings via PyO3 and can be linked directly into custom Rust applications by adding it to Cargo.toml.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →