# What Is the Difference Between switchyard-server and switchyard launch Command?

> Understand the difference between switchyard-server and switchyard launch. Learn how the Rust binary hosts the LLM-proxy API and the Python CLI orchestrates coding agents for tools like Claude Code.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-17

---

**The `switchyard-server` is a standalone Rust binary that hosts the core LLM-proxy HTTP API, while `switchyard launch` is a Python CLI command that orchestrates coding agents by starting the Rust server in-process and forwarding control to tools like Claude Code or Codex CLI.**

The NVIDIA-NeMo/Switchyard repository provides two distinct entry points for running the LLM routing platform. Understanding the difference between switchyard-server and switchyard launch command is essential for choosing the right deployment strategy for production services versus local development workflows.

## Architectural Overview

### switchyard-server (Native Rust Binary)

`switchyard-server` is a native **Rust** binary located in [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs). It runs independently of Python and hosts TOML-defined deployments that directly expose the LLM-proxy HTTP API, including endpoints like `/v1/chat/completions`, `/v1/models`, and `/v1/stats`. This binary handles routing logic, algorithm execution, observability, and metrics collection through a standalone process that responds to OS signals (`SIGINT`, `SIGTERM`) for lifecycle management.

### switchyard launch (Python CLI Wrapper)

The `switchyard launch` command is implemented in **Python** within [`switchyard/cli/switchyard_cli.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/switchyard_cli.py) and [`switchyard/cli/launch_command.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launch_command.py). It provides a high-level entry point for developers to run coding agents (Claude Code, Codex CLI, OpenClaw) against a Switchyard deployment. Rather than running independently, this command creates a transient, embedded server instance through PyO3 bindings and automatically tears it down when the agent exits.

## Key Differences in Implementation

The distinction between these two components spans implementation language, lifecycle management, and intended use cases.

| Aspect | `switchyard-server` | `switchyard launch` |
|--------|---------------------|---------------------|
| **Language** | Native Rust (`crates/switchyard-server`) | Python (`switchyard/cli/`) with PyO3 bindings |
| ** Execution** | Standalone binary: `switchyard-server --config routes.toml` | CLI subcommand: `switchyard launch claude --config routes.toml` |
| **Process Lifecycle** | Long-running daemon; manual shutdown via signals | Transient; starts `NativeServer` in-process and stops on exit |
| **API Exposure** | Direct HTTP server on configured port | Internal server accessed by the launched agent |
| **Primary Use** | Production deployments, external API consumers | Local development, quick agent testing |

When you invoke `switchyard launch`, the Python code in [`switchyard/cli/launchers/native_server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/native_server.py) constructs a `switchyard_rust.server.Server` instance via PyO3, forwards the agent-specific arguments, and returns the agent's exit status to the shell.

## Practical Usage Examples

### Running the Native Server Directly

For production deployments or when exposing the LLM proxy to external services, execute the Rust binary directly:

```bash

# Start a Switchyard deployment defined by routes.toml

switchyard-server --config routes.toml

```

This command reads the configuration from [`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs) and initializes the routing layer defined in [`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs). The server continues running until it receives a termination signal, making it suitable for containerized deployments or systemd services.

### Launching a Coding Agent via CLI

For development workflows that require immediate interaction with Claude Code or similar tools, use the Python wrapper:

```bash

# Launch Claude Code against the deployment

switchyard launch claude \
    --model switchyard \
    --config routes.toml \
    -- claude-args --some-flag

```

This invocation triggers the logic in [`switchyard/cli/launch_command.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launch_command.py), which parses the subcommand arguments, instantiates the `NativeServer` class from [`switchyard/cli/launchers/native_server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/native_server.py), and delegates to the appropriate agent launcher (e.g., [`switchyard/cli/launchers/claude_code_launcher.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/claude_code_launcher.py)).

### Programmatic Server Management (Advanced)

You can also embed the server within Python applications using the `NativeServer` wrapper:

```python
from switchyard.cli.launchers.native_server import NativeServer
from pathlib import Path

# Create a server instance programmatically

server = NativeServer(config=Path("routes.toml"))
print(f"Server listening at {server.base_url}")

# The server runs until explicitly closed or the Python process exits

server.close()

```

## Source Code Structure and Key Files

Understanding the file organization clarifies the separation of concerns between the Rust core and Python tooling.

- **[`crates/switchyard-server/src/main.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/main.rs)** – Binary entry point for the native Rust server
- **[`crates/switchyard-server/src/lib.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-server/src/lib.rs)** – Core server implementation, routing logic, and metrics collection
- **[`switchyard/cli/switchyard_cli.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/switchyard_cli.py)** – Top-level Click or argparse definition for the `switchyard` command
- **[`switchyard/cli/launch_command.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launch_command.py)** – Implementation of the `launch` subcommand logic
- **[`switchyard/cli/launchers/native_server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/native_server.py)** – Python wrapper that manages the Rust server lifecycle via PyO3
- **[`switchyard/cli/launchers/claude_code_launcher.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/claude_code_launcher.py)** – Agent-specific launch logic for Claude Code integration

## Summary

- **`switchyard-server`** is the high-performance Rust binary that provides the core LLM-proxy HTTP API for production deployments
- **`switchyard launch`** is a convenience-focused Python CLI command that starts the Rust server in-process, runs a coding agent, and handles cleanup automatically
- Use the **native binary** when deploying standalone routing services or integrating with external infrastructure
- Use the **launch command** for local development workflows requiring quick iteration with AI coding assistants
- The Python wrapper in [`switchyard/cli/launchers/native_server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/native_server.py) bridges the two worlds by embedding the Rust server via PyO3 bindings

## Frequently Asked Questions

### Can I run switchyard launch in production environments?

While possible, `switchyard launch` is designed for **development workflows** rather than production. It creates a transient server that terminates when the agent exits, which is unsuitable for long-running API services. For production, deploy `switchyard-server` directly as a standalone binary or containerized service.

### What happens to the server when the launch command exits?

The server is **automatically terminated**. The `NativeServer` class in [`switchyard/cli/launchers/native_server.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard/cli/launchers/native_server.py) manages the Rust server's lifecycle as a context manager; when the Python process exits or the context ends, it signals the embedded `switchyard_rust.server.Server` instance to shut down gracefully.

### Do I need Python installed to run switchyard-server?

**No.** The `switchyard-server` binary is a compiled Rust artifact that runs independently of Python. You only need Python when using the `switchyard launch` command or other CLI utilities defined in the `switchyard/cli/` directory.

### Which command should I use for running Claude Code against Switchyard?

Use **`switchyard launch claude`**. This command handles the boilerplate of starting the server, ensuring the correct model routing is active, and forwarding all subsequent arguments to the Claude Code CLI. It eliminates the need to manually manage two separate processes (server and agent) in different terminals.