# Main Directories in the NVIDIA-NeMo Switchyard Repository: A Complete Guide

> Explore the main directories of the NVIDIA-NeMo Switchyard repository. Understand the Python API, Rust core, tests, examples, and more for efficient LLM orchestration.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-23

---

**The Switchyard repository organizes its LLM orchestration stack into ten primary directories, spanning from the public Python API in `switchyard/` and `switchyard_rust/` to the high-performance Rust core in `crates/`, with supporting infrastructure in `tests/`, `examples/`, `benchmark/`, and CI tooling.**

Switchyard is a Python-centric orchestration layer that routes LLM traffic to multiple back-ends while providing A/B testing, health-aware routing, and format translation. Understanding the main directories in the Switchyard repository is essential for developers extending its routing algorithms or deploying the native server. The codebase follows a hybrid Python-Rust architecture where Python provides the ergonomic API surface and Rust delivers the high-performance core.

## Core Python Packages

### `switchyard/` – Public Python API

The `switchyard/` directory houses the public Python package that most developers interact with daily. According to the source code, this directory exposes version information, type-checked wrappers, and async client helpers through [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py).

This package provides the `SwitchyardClient` class used for asynchronous LLM inference:

```python
import asyncio
from switchyard import SwitchyardClient

async def main():
    # The client automatically discovers the local server if `SWITCHYARD_URL` is set.

    client = SwitchyardClient()
    response = await client.chat(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "What is Switchyard?"}]
    )
    print("LLM answer:", response.choices[0].message["content"])

if __name__ == "__main__":
    asyncio.run(main())

```

### `switchyard_rust/` – PyO3 Bridge Facade

The `switchyard_rust/` directory contains a thin Python façade that loads the compiled native extension. As implemented in [`switchyard_rust/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/__init__.py), this package bridges the gap between Python code and the Rust binary, enabling zero-cost access to the core routing logic from Python scripts.

Developers can access low-level protocol types directly through this bridge:

```python
import switchyard_rust as srt

# Access low‑level protocol types defined in the Rust crate

request = srt.ProtocolRequest(
    model="gpt-4o-mini",
    messages=[srt.Message(role="user", content="Hello")]
)
response = srt.send_request(request)   # Calls into the compiled Rust lib

print(response)

```

## Native Implementation

### `crates/` – Rust Core Crates

The `crates/` directory contains all Rust crates that constitute the performance-critical core of Switchyard. As defined in the root [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) workspace, these crates include routing algorithms, protocol types, the HTTP server, translation codecs, and the PyO3 bridge.

Key crates include:

- **`switchyard-server`** – The native HTTP server implementing OpenAI and Anthropic API compatibility, with its entry point at [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs)
- **`switchyard-translation`** – Handles wire-format translation between provider-neutral protocols and provider-specific JSON payloads
- **`switchyard-py`** – Contains the PyO3 glue code exposing Rust structs to Python, located in [`crates/switchyard-py/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-py/src/lib.rs)

To start the native server directly:

```bash

# Install the compiled wheel (or use uv to install from source)

uv pip install .

# Run the server with a minimal TOML configuration

switchyard-server --config examples/prometheus/switchyard.rules.yaml --port 4000

```

## Validation and Examples

### `tests/` – Integration Test Suite

The `tests/` directory contains the Pytest test suite covering the Python layer, the Rust-Python bridge, and end-to-end integration scenarios. The file [`tests/test_libsy_minimal_bindings.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_libsy_minimal_bindings.py) specifically guarantees that the Python-Rust bindings remain functional across releases, validating that changes to `crates/switchyard-py/` do not break the public API.

### `examples/` – Usage Patterns

This directory provides ready-to-run scripts demonstrating common usage patterns, including basic client initialization, Prometheus metrics export, and custom routing algorithm implementation. The file [`examples/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/examples/libsy.py) demonstrates building a custom routing algorithm with the `libsy` Python wrapper.

### `benchmark/` – Performance Measurement

The `benchmark/` directory contains tools for measuring routing latency, sample server configurations, and dataset preparation scripts. The [`benchmark/run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run_manifest.py) file provides a CLI entry point for performance profiling.

To run benchmarks programmatically:

```python

# benchmark/run_manifest.py provides a CLI entry point; here we invoke programmatically

from benchmark.run_manifest import run_profile

await run_profile(
    config_path="benchmark/server-configs/tb-lite-single-gpt-5-5.toml",
    profile_name="tb-lite-gpt",
    duration_seconds=30
)

```

### `dev-server/` – Development Configuration

The `dev-server/` directory holds development-time server configuration, including a systemd service file (`dev-server/switchyard.service`), example TOML configurations, and deployment documentation.

## Repository Infrastructure

### `.github/` – Continuous Integration

The `.github/` directory stores GitHub Actions workflows, issue and PR templates, and code ownership definitions. The file [`.github/workflows/ci.yml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/.github/workflows/ci.yml) defines the CI pipeline that runs both `cargo test` and `pytest` on every pull request, ensuring cross-language compatibility.

### `.agents/` – Automation Skills

This repository-internal directory contains "skills" that automate CI actions, such as Rust code review and testing workflows. The file [`.agents/skills/switchyard-testing-ci/SKILL.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/.agents/skills/switchyard-testing-ci/SKILL.md) documents automated testing procedures.

### `assets/` – Static Resources

The `assets/` directory contains static resources such as the project logo (`assets/logo.png`).

## Summary

- **`switchyard/`** exposes the high-level async Python API (`SwitchyardClient`) through [`switchyard/__init__.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/__init__.py)
- **`switchyard_rust/`** provides the PyO3 bridge loading the compiled Rust extension
- **`crates/`** contains the Rust workspace with server, translation, and binding crates
- **`tests/`**, **`examples/`**, and **`benchmark/`** provide validation, documentation, and performance tooling
- **`.github/`** and **`.agents/`** manage CI automation and repository governance
- Root files [`pyproject.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/pyproject.toml) and [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) define the Python and Rust build metadata respectively

## Frequently Asked Questions

### What is the difference between the `switchyard/` and `switchyard_rust/` directories?

The `switchyard/` directory contains the idiomatic Python API with type hints and async helpers, while `switchyard_rust/` serves as a thin loader for the native PyO3 extension compiled from `crates/switchyard-py/`. Most applications should import from `switchyard`, but advanced users requiring direct access to Rust protocol types can use `switchyard_rust`.

### Where is the main server binary defined in the repository?

The native HTTP server binary is defined in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs), with its package configuration in [`crates/switchyard-server/Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/Cargo.toml). This server implements the OpenAI-compatible API surface and handles health-aware routing in Rust.

### How does the repository handle testing across both Python and Rust?

The `tests/` directory contains Pytest suites that validate the Python-Rust bridge, while the root [`Cargo.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/Cargo.toml) workspace enables `cargo test` for native Rust unit tests. The [`.github/workflows/ci.yml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/.github/workflows/ci.yml) orchestrates both test suites in CI to prevent regressions in the PyO3 bindings.

### Which directory contains configuration files for running benchmarks?

Benchmark configurations reside in `benchmark/server-configs/`, while the execution harness is located at [`benchmark/run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run_manifest.py). This Python module provides both CLI and programmatic interfaces for latency profiling against specific routing profiles.