# How Is the Switchyard Codebase Structured? Inside NVIDIA-NeMo's Python-Rust Hybrid Architecture

> Explore the NVIDIA NeMo Switchyard codebase structure, a Python Rust hybrid architecture. Discover how it separates API ergonomics from high-performance logic for efficient routing and translation.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-08-23

---

**Switchyard is organized as a mixed-language project combining a Python integration layer with a native Rust implementation, separating public API ergonomics from high-performance routing and translation logic.**

The Switchyard codebase structure follows a strict separation of concerns between user-facing interfaces and systems-level performance. As a hybrid Python-Rust project maintained by NVIDIA under the NeMo ecosystem, Switchyard enables intelligent LLM routing while maintaining clean abstractions. Understanding its directory layout and architectural flow is essential for extending the router or embedding it in production pipelines.

## High-Level Directory Layout

Switchyard partitions its functionality into distinct directories that isolate language-specific concerns. The repository follows this organizational hierarchy:

| Layer | Directory | Purpose |
|-------|-----------|---------|
| **Python Public API** | `switchyard/` | Exposes the top-level package and thin wrappers that forward calls to the Rust backend. |
| **Python-Rust Bridge** | `switchyard_rust/` | Contains PyO3 bindings and high-level Python wrappers that load the compiled Rust shared library. |
| **Rust Core** | `crates/` | Implements routing algorithms, protocol types, translation codecs, and the standalone server. |
| **Documentation** | `docs/` | Markdown files describing architecture, usage patterns, and CLI references. |
| **Testing** | `tests/` and `tests/e2e/` | Unit tests for both languages plus end-to-end integration tests. |
| **Examples & Scripts** | `examples/` and `scripts/` | Sample client code and benchmarking utilities. |

Build configuration is split between **Python** ([`pyproject.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/pyproject.toml)) and **Rust** ([`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml)) manifests, with a `Dockerfile` supporting containerized deployments.

## The Python Public API Layer

The `switchyard/` directory provides the primary entry point for developers. In [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py), the package exports high-level client classes like `ChatClient` that implement OpenAI-compatible interfaces.

The Python layer focuses on ergonomics rather than heavy computation. In [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py), thin Python wrappers expose routing primitives to the public API while delegating actual execution to Rust. This design allows Python code to handle configuration loading and request marshalling while the Rust core manages computational routing logic.

## The Rust Core and Python-Rust Bridge

The `switchyard_rust/` directory houses the **PyO3 bindings** that bridge the two languages. The file [`switchyard_rust/_native.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/_native.py) serves as the low-level loader for the compiled Rust extension (`.so` or `.pyd` file), handling dynamic library initialization.

High-level abstractions reside in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), which maps Python method calls to Rust primitives defined in the `crates/libsy` workspace. The standalone server implementation lives in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py), where an Axum-based HTTP server exposes REST endpoints that implement OpenAI and Anthropic API shapes.

The underlying Rust implementation in `crates/` (implied by the workspace configuration) contains the actual routing engine, protocol normalization logic, and translation codecs that convert between provider-specific request formats.

## Request Lifecycle and Data Flow

Understanding the Switchyard codebase structure requires tracing the five-phase request lifecycle documented in [`docs/architecture.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/architecture.md):

1. **Client Interaction** — Users import from `switchyard` and invoke methods like `ChatClient.chat()`. The Python layer in [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py) handles initial request validation.

2. **Normalization** — Incoming OpenAI or Anthropic-style requests cross into Rust via the PyO3 boundary. The Rust `protocol` crate converts these into provider-neutral request objects.

3. **Routing** — The core routing algorithm (implemented in `crates/libsy`) applies policies such as weighted distribution or classifier-based selection. Python wrappers in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) expose these strategies.

4. **Execution** — Selected backends receive requests through the translation layer (`switchyard-translation` crate), which manages wire-format conversion.

5. **Response** — Results stream back through the Rust-Python boundary, preserving the original client API shape while leveraging Rust's async runtime for performance.

## Practical Usage Patterns

### Creating a Chat Client

To interact with the routing system, instantiate the client through the public API:

```python
import switchyard as sy

# Load default configuration (expects a routes.toml file)

client = sy.ChatClient()

# Send a chat request using OpenAI-compatible format

response = client.chat(
    messages=[{"role": "user", "content": "Hello, Switchyard!"}]
)

print(response.choices[0].message.content)

```

The `ChatClient` class defined in [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py) ultimately delegates to the Rust implementation through the bridge layer.

### Defining Custom Routing Policies

For advanced use cases, access the Rust primitives directly through the bridge:

```python
import switchyard_rust as sry

# Create a weighted router

router = sry.WeightedRouter(
    routes=[
        sry.Route(target="openai", weight=70),
        sry.Route(target="nvidia", weight=30),
    ]
)

# Attach to a client instance

client = sry.ChatClient(router=router)

```

The `WeightedRouter` class is implemented in the Rust `libsy` crate and exposed via [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).

### Running the Standalone Server

Deploy the Rust-based HTTP server using the CLI entry point:

```bash

# From repository root

switchyard-server --config routes.toml --port 4000

```

This command invokes the Axum server wired in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py), which bypasses Python interpreter overhead for maximum throughput.

## Summary

- **Switchyard** separates concerns into `switchyard/` (Python API) and `switchyard_rust/` (PyO3 bridge) with the heavy lifting performed by Rust crates.
- **Key entry points** include [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py) for client access and [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py) for standalone deployment.
- **Routing logic** resides in the Rust workspace but is accessible through Python wrappers in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py).
- **Request flow** moves from Python normalization through Rust routing and translation before returning provider-specific responses.
- **Build system** uses standard Python packaging ([`pyproject.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/pyproject.toml)) alongside Cargo workspaces ([`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml)) to manage the mixed-language dependencies.

## Frequently Asked Questions

### What is the difference between the `switchyard` and `switchyard_rust` directories?

The `switchyard/` directory contains the public Python API intended for end users, exposing high-level classes like `ChatClient`. The `switchyard_rust/` directory contains the **PyO3 bridge** code—specifically [`_native.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/_native.py) for low-level extension loading and [`libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/libsy.py) for high-level Rust bindings. While `switchyard/` focuses on ergonomics and configuration, `switchyard_rust/` handles the FFI boundary to the compiled Rust shared library.

### How does Python code communicate with the Rust backend?

Communication occurs through **PyO3 bindings** defined in [`switchyard_rust/_native.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/_native.py), which loads the compiled Rust extension. High-level wrappers in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) expose Rust structs and methods as Python classes. When you call a method on `ChatClient`, the Python layer serializes arguments, crosses the FFI boundary into Rust, executes the routing logic in the `crates/` workspace, and deserializes the results back into Python objects.

### Where are the routing algorithms actually implemented?

The core routing algorithms are implemented in **Rust** within the `crates/libsy` crate (referenced by the workspace configuration in [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml)). However, Python wrappers in [`switchyard/libsy/algorithms.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/libsy/algorithms.py) provide idiomatic interfaces to these algorithms. For example, `WeightedRouter` logic runs in compiled Rust code, but you instantiate and configure it through the Python API in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py).

### How do I run Switchyard as a standalone server without writing Python code?

You can launch the native HTTP server using the command-line interface. The server implementation in [`switchyard_rust/server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/server.py) initializes a Rust Axum server that listens for OpenAI-compatible HTTP requests. Run `switchyard-server --config routes.toml --port 4000` from your shell to start the server, which operates independently of the Python interpreter after initialization.