What Is the Difference Between switchyard-server and switchyard launch Command?
The switchyard-server is a standalone Rust binary that hosts the core LLM-proxy HTTP API, while switchyard launch is a Python CLI command that orchestrates coding agents by starting the Rust server in-process and forwarding control to tools like Claude Code or Codex CLI.
The NVIDIA-NeMo/Switchyard repository provides two distinct entry points for running the LLM routing platform. Understanding the difference between switchyard-server and switchyard launch command is essential for choosing the right deployment strategy for production services versus local development workflows.
Architectural Overview
switchyard-server (Native Rust Binary)
switchyard-server is a native Rust binary located in crates/switchyard-server/src/main.rs. It runs independently of Python and hosts TOML-defined deployments that directly expose the LLM-proxy HTTP API, including endpoints like /v1/chat/completions, /v1/models, and /v1/stats. This binary handles routing logic, algorithm execution, observability, and metrics collection through a standalone process that responds to OS signals (SIGINT, SIGTERM) for lifecycle management.
switchyard launch (Python CLI Wrapper)
The switchyard launch command is implemented in Python within switchyard/cli/switchyard_cli.py and switchyard/cli/launch_command.py. It provides a high-level entry point for developers to run coding agents (Claude Code, Codex CLI, OpenClaw) against a Switchyard deployment. Rather than running independently, this command creates a transient, embedded server instance through PyO3 bindings and automatically tears it down when the agent exits.
Key Differences in Implementation
The distinction between these two components spans implementation language, lifecycle management, and intended use cases.
| Aspect | switchyard-server |
switchyard launch |
|---|---|---|
| Language | Native Rust (crates/switchyard-server) |
Python (switchyard/cli/) with PyO3 bindings |
| ** Execution** | Standalone binary: switchyard-server --config routes.toml |
CLI subcommand: switchyard launch claude --config routes.toml |
| Process Lifecycle | Long-running daemon; manual shutdown via signals | Transient; starts NativeServer in-process and stops on exit |
| API Exposure | Direct HTTP server on configured port | Internal server accessed by the launched agent |
| Primary Use | Production deployments, external API consumers | Local development, quick agent testing |
When you invoke switchyard launch, the Python code in switchyard/cli/launchers/native_server.py constructs a switchyard_rust.server.Server instance via PyO3, forwards the agent-specific arguments, and returns the agent's exit status to the shell.
Practical Usage Examples
Running the Native Server Directly
For production deployments or when exposing the LLM proxy to external services, execute the Rust binary directly:
# Start a Switchyard deployment defined by routes.toml
switchyard-server --config routes.toml
This command reads the configuration from crates/switchyard-server/src/main.rs and initializes the routing layer defined in crates/switchyard-server/src/lib.rs. The server continues running until it receives a termination signal, making it suitable for containerized deployments or systemd services.
Launching a Coding Agent via CLI
For development workflows that require immediate interaction with Claude Code or similar tools, use the Python wrapper:
# Launch Claude Code against the deployment
switchyard launch claude \
--model switchyard \
--config routes.toml \
-- claude-args --some-flag
This invocation triggers the logic in switchyard/cli/launch_command.py, which parses the subcommand arguments, instantiates the NativeServer class from switchyard/cli/launchers/native_server.py, and delegates to the appropriate agent launcher (e.g., switchyard/cli/launchers/claude_code_launcher.py).
Programmatic Server Management (Advanced)
You can also embed the server within Python applications using the NativeServer wrapper:
from switchyard.cli.launchers.native_server import NativeServer
from pathlib import Path
# Create a server instance programmatically
server = NativeServer(config=Path("routes.toml"))
print(f"Server listening at {server.base_url}")
# The server runs until explicitly closed or the Python process exits
server.close()
Source Code Structure and Key Files
Understanding the file organization clarifies the separation of concerns between the Rust core and Python tooling.
crates/switchyard-server/src/main.rs– Binary entry point for the native Rust servercrates/switchyard-server/src/lib.rs– Core server implementation, routing logic, and metrics collectionswitchyard/cli/switchyard_cli.py– Top-level Click or argparse definition for theswitchyardcommandswitchyard/cli/launch_command.py– Implementation of thelaunchsubcommand logicswitchyard/cli/launchers/native_server.py– Python wrapper that manages the Rust server lifecycle via PyO3switchyard/cli/launchers/claude_code_launcher.py– Agent-specific launch logic for Claude Code integration
Summary
switchyard-serveris the high-performance Rust binary that provides the core LLM-proxy HTTP API for production deploymentsswitchyard launchis a convenience-focused Python CLI command that starts the Rust server in-process, runs a coding agent, and handles cleanup automatically- Use the native binary when deploying standalone routing services or integrating with external infrastructure
- Use the launch command for local development workflows requiring quick iteration with AI coding assistants
- The Python wrapper in
switchyard/cli/launchers/native_server.pybridges the two worlds by embedding the Rust server via PyO3 bindings
Frequently Asked Questions
Can I run switchyard launch in production environments?
While possible, switchyard launch is designed for development workflows rather than production. It creates a transient server that terminates when the agent exits, which is unsuitable for long-running API services. For production, deploy switchyard-server directly as a standalone binary or containerized service.
What happens to the server when the launch command exits?
The server is automatically terminated. The NativeServer class in switchyard/cli/launchers/native_server.py manages the Rust server's lifecycle as a context manager; when the Python process exits or the context ends, it signals the embedded switchyard_rust.server.Server instance to shut down gracefully.
Do I need Python installed to run switchyard-server?
No. The switchyard-server binary is a compiled Rust artifact that runs independently of Python. You only need Python when using the switchyard launch command or other CLI utilities defined in the switchyard/cli/ directory.
Which command should I use for running Claude Code against Switchyard?
Use switchyard launch claude. This command handles the boilerplate of starting the server, ensuring the correct model routing is active, and forwarding all subsequent arguments to the Claude Code CLI. It eliminates the need to manually manage two separate processes (server and agent) in different terminals.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →