How to Configure the Server Shutdown Timeout for Graceful Draining in NVIDIA Switchyard

Switchyard performs a graceful drain on shutdown by waiting for in-flight requests to complete, with a default timeout of 30 seconds that you can override via the --shutdown_timeout CLI flag, the ServerRunOptions struct in Rust, or the timeout_secs parameter in Python.

NVIDIA Switchyard is a high-performance HTTP inference server that supports graceful shutdown to prevent abrupt termination of active requests. Configuring the server shutdown timeout ensures your deployment waits long enough for in-flight inference to complete before exiting. This guide covers three methods to set this timeout using the official NVIDIA-NeMo/Switchyard source code.

Understanding the Graceful Shutdown Timeout

The graceful shutdown mechanism stops accepting new connections while allowing existing requests to finish processing. According to the source code in crates/switchyard-server/src/lib.rs, the constant DEFAULT_GRACEFUL_SHUTDOWN_TIMEOUT defines a 30-second default duration. When the shutdown signal triggers, the server invokes handle.graceful_shutdown(Some(timeout)) within the serve_until_shutdown function to begin the draining process.

Configuring the Shutdown Timeout

Command-Line Interface (Rust Binary)

When running the switchyard-server binary, pass the --shutdown_timeout flag with a human-readable duration format. The CLI parser in crates/switchyard-server/src/cli.rs (lines 43-46) populates the ServerRunOptions.shutdown_timeout field with your specified value.


# Run the server with a 45-second graceful shutdown window

switchyard-server \
  --config path/to/deployment.toml \
  --shutdown_timeout 45s

Programmatic Rust API

For embedded deployments, construct a ServerRunOptions struct and set the shutdown_timeout field to a std::time::Duration before calling run_server. The function signature in crates/switchyard-server/src/lib.rs (lines 90-101) accepts these options to configure the server's behavior.

use switchyard_server::{run_server, ServerRunOptions, ServerState};

let (state, _) = load_server_state("deployment.toml")?;
let options = ServerRunOptions {
    addr: "0.0.0.0:4000".parse().unwrap(),
    backlog: switchyard_server::DEFAULT_LISTEN_BACKLOG,
    dry_run: false,
    shutdown_timeout: std::time::Duration::from_secs(60), // 1-minute timeout
    tls: None,
};

run_server(state, options).await?;

Python Bindings

When using switchyard-py, call Server.close(timeout_secs=...) where the parameter accepts a float representing seconds. Note that the Python default differs from the Rust server: the constant DEFAULT_SHUTDOWN_TIMEOUT_SECS in crates/switchyard-py/src/server_bindings.rs (lines 22-30) sets a 2-second default for the Python API.

from switchyard_rust.server import Server

# Start a loopback server

srv = Server("deployment.toml")
print(f"Listening at {srv.base_url}")

# Later, request a graceful shutdown with a 5-second timeout

srv.close(timeout_secs=5.0)

Implementation Details and Timeout Behavior

Internally, Switchyard converts your timeout parameter into a std::time::Duration. If active requests complete before the timeout expires, the server shuts down immediately. However, if the timeout elapses first, the server aborts remaining tasks. This logic resides in close_inner within crates/switchyard-py/src/server_bindings.rs (lines 33-51), which handles the forced termination of dangling connections.

Summary

  • The default graceful shutdown timeout is 30 seconds in the Rust server but 2 seconds in Python bindings.
  • Override via CLI with --shutdown_timeout, Rust with ServerRunOptions.shutdown_timeout, or Python with timeout_secs.
  • The server calls graceful_shutdown in serve_until_shutdown to initiate the draining process.
  • Unfinished requests are aborted if they exceed the configured timeout during shutdown.

Frequently Asked Questions

What happens if requests exceed the shutdown timeout?

If in-flight requests do not complete before the timeout expires, Switchyard aborts the remaining tasks and proceeds with shutdown. This prevents the server from hanging indefinitely during deployment rollouts or pod terminations in Kubernetes.

Can I disable the shutdown timeout entirely?

While the internal graceful_shutdown method accepts None to wait indefinitely, the current Switchyard implementations require a finite duration. You can effectively disable the timeout by setting an extremely large value, such as 24h in the CLI or Duration::from_secs(86400) in Rust.

Why does the Python API use a different default timeout?

The Python bindings default to 2 seconds (DEFAULT_SHUTDOWN_TIMEOUT_SECS) because Python-based inference scripts often run in interactive or testing environments where rapid shutdown is preferred. The Rust server assumes production workloads requiring the standard 30-second window.

How do I format the duration for the CLI flag?

The --shutdown_timeout flag accepts human-readable formats such as 30s for 30 seconds, 2m for 2 minutes, or 1h for 1 hour. The parser converts these strings into a Duration before passing them to the server configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →