System Requirements for NVIDIA NeMo Switchyard: Hardware, Software, and Build Prerequisites
To install and run Switchyard, you need a Linux system with an x86_64-v3/AVX2 or Neoverse N1 CPU, Python 3.10+ or Rust 1.96.1+, and standard development tools including cargo or optionally uv.
NVIDIA NeMo Switchyard is a multi-language LLM routing library that supports Python embedding, Rust integration, and standalone HTTP proxy deployment. Understanding the system requirements for Switchyard ensures compatibility whether you are installing the nemo-switchyard package from PyPI or compiling the switchyard-server binary from the NVIDIA-NeMo/Switchyard source. The following specifications are derived directly from the INSTALLATION.md and build configuration files in the repository.
Hardware Prerequisites
CPU Architecture Requirements
Switchyard distributes pre-built wheels for two specific CPU classes. According to INSTALLATION.md, your processor must meet one of the following specifications:
- x86_64 systems: Requires x86_64-v3 microarchitecture level with AVX2 instruction set support. Run
lscpu | grep avx2to verify your CPU flags. - ARM64 systems: Requires Neoverse N1 class CPU or equivalent. This applies to the aarch64 wheels distributed for the package.
If your CPU does not meet these requirements, you must compile Switchyard from source using the Rust toolchain instead of using the pre-built wheels.
GPU and Accelerator Requirements
Switchyard itself does not require a GPU. The library functions purely as a routing layer that directs requests to downstream model endpoints (such as OpenRouter, OpenAI, or Anthropic). GPU requirements only apply to the local inference servers you might configure as targets, not to the Switchyard routing component itself.
Software Dependencies
Operating System
Switchyard requires Linux (any modern distribution). The project builds and tests against recent kernels (≥ 4.15). While the Rust components may compile on other platforms, the official INSTALLATION.md only documents Linux support for both x86_64 and aarch64 architectures.
Python Runtime Requirements
For the nemo-switchyard package, you need:
- Python 3.10 or newer
- pip or uv (the latter recommended for development builds)
The Python bindings have no additional runtime dependencies beyond the native Rust extension. The entry point for these bindings is located in crates/switchyard-py/src/lib.rs, which exposes the core Rust library to Python.
Rust Toolchain Requirements
For building the switchyard-server binary, Rust libraries (switchyard-libsy, switchyard-protocol, etc.), or the NeMo Relay plugin, you need:
- Rust 1.96.1 or newer
- cargo (Rust's package manager)
The workspace definition in Cargo.toml at the repository root enforces these version constraints and manages the crate metadata for the multi-crate project.
Build Tools and System Packages
When compiling from source rather than using pre-built wheels, install standard Linux development tools:
sudo apt-get install gcc make libssl-dev pkg-config
For Python development workflows, the repository recommends installing uv, a fast Python package manager:
pip install uv
This tool is referenced in the development section of INSTALLATION.md for managing Python dependencies during local builds.
Installation Methods by Use Case
Installing Python Bindings
To embed Switchyard in Python applications, install the wheel from PyPI. This requires Python 3.10+ but does not require a Rust toolchain unless you are building from source:
pip install nemo-switchyard
The following example from examples/libsy.py demonstrates a minimal embedding using the stage_router algorithm from crates/libsy/src/algorithms.rs:
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router
# Create a stage-router that prefers the "efficient" model
algorithm = stage_router(
"capable",
"efficient",
picker="efficient_first",
confidence_threshold=0.5,
)
async def call_with_fallback(request, models, clients):
for model in models:
try:
return LlmResponse.Agg(await clients[model].call(request))
except Exception:
continue
raise RuntimeError("All candidates failed")
async def route(request: dict, clients: dict):
async for step in algorithm.run_stream(request):
match step:
case Step.CallModel(call):
call.respond(await call_with_fallback(call.request, call.models, clients))
case Step.Done(outcome):
return outcome.response or await call_with_fallback(
outcome.request, outcome.selected_model_ids, clients
)
Building the Standalone Rust Server
To run Switchyard as an HTTP proxy, install the binary using cargo (Rust 1.96.1+ required):
cargo install --locked switchyard-server
Create a configuration file following the TOML schema documented in docs/reference/toml_schema.md:
schema_version = 1
[llm_clients.openrouter]
format = "openai_chat"
base_url = "https://openrouter.ai/api/v1"
api_key_env = "OPENROUTER_API_KEY"
[targets.capable]
id = "anthropic/claude-opus-4.8"
llm_client = "openrouter"
[targets.efficient]
id = "z-ai/glm-5.2"
llm_client = "openrouter"
[routes.switchyard]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
Run the server with:
export OPENROUTER_API_KEY="your-key"
switchyard-server --config routes.toml --host 127.0.0.1 --port 4000
Compiling the NeMo Relay Plugin
To use Switchyard as a plugin for NeMo Relay (requires NeMo Relay ≥ 0.8.1), compile from the repository root:
git clone https://github.com/NVIDIA-NeMo/Switchyard.git
cd Switchyard
cargo build --release -p switchyard-nemo-relay-plugin
Copy the generated relay-plugin.toml to your Relay plugin directory and configure the switchyard_config_path as detailed in crates/switchyard-nemo-relay-plugin/README.md.
Network and Configuration Requirements
Switchyard requires outbound HTTP/HTTPS access to whatever model providers you configure (e.g., OpenRouter, OpenAI, Anthropic). Ensure your firewall allows connections to these endpoints on standard web ports.
For complete configuration options, refer to the TOML schema reference in docs/reference/toml_schema.md and the server documentation in crates/switchyard-server/README.md.
Summary
- Hardware: Linux with x86_64-v3/AVX2 or Neoverse N1 CPU; no GPU required for the router itself.
- Software: Python 3.10+ for Python bindings, or Rust 1.96.1+ for server and library compilation.
- Build Tools:
cargois mandatory for Rust builds;gcc,make, andlibssl-devfor source compilation;uvoptional for Python development. - Network: Outbound HTTPS to configured LLM providers.
Frequently Asked Questions
Does Switchyard require a GPU to run?
No. Switchyard operates as a routing layer and does not perform model inference itself. According to the source code architecture, it simply routes requests to downstream endpoints. Any GPU requirements belong to those downstream inference servers, not to the Switchyard installation.
Can I run Switchyard on macOS or Windows?
The official INSTALLATION.md only documents Linux support for both x86_64 and aarch64 architectures. While the Rust code may compile on other platforms, the pre-built wheels target Linux specifically, and the project does not guarantee functionality on macOS or Windows.
What is the difference between Python and Rust installation requirements?
Python installation requires only the runtime (3.10+) and uses pre-compiled wheels containing native extensions. Rust installation requires the full toolchain (1.96.1+) including cargo to compile the switchyard-server binary or to build the library from source when your CPU does not support the pre-built wheel requirements (AVX2 or Neoverse N1).
Why does Switchyard require specific CPU instruction sets?
The pre-built wheels in INSTALLATION.md target x86_64-v3/AVX2 for x86 systems and Neoverse N1 for ARM64 systems to ensure optimal performance for the routing algorithms implemented in crates/libsy/src/algorithms.rs. These instruction sets provide the necessary SIMD capabilities for efficient request processing and classification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →