# Switchyard | NVIDIA-NeMo | Knowledge Base | Instagit

Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.

GitHub Stars: 1.3k

Repository: https://github.com/NVIDIA-NeMo/Switchyard

---

## Articles

### [How Routing Overhead Is Measured During Switchyard Soak Testing](/NVIDIA-NeMo/Switchyard/how-is-routing-overhead-measured-during-switchyard-soak-testing-and-how-does-this-measurement-differ-from-end-to-end-latency)

Learn how Switchyard isolates routing overhead during soak testing. Discover how measuring latency delta reveals Switchyard specific processing time vs direct arm latency.

- Tags: performance
- Published: 2026-09-13

### [How the closed_book_proxy Integration in Switchyard Uses an Allowlist Rewriter to Restrict Agent Traffic During Benchmarking](/NVIDIA-NeMo/Switchyard/how-does-the-closed-book-proxy-integration-in-switchyard-utilize-an-allowlist-rewriter-to-restrict-agent-traffic-during-benchmarking)

Learn how Switchyard's closed_book_proxy uses an allowlist rewriter with Mitmproxy to block unauthorized agent traffic during benchmarking, ensuring data integrity and efficient testing.

- Tags: how-to-guide
- Published: 2026-09-13

### [How to Reproduce Terminal-Bench 2.1 Benchmark Results for Switchyard and Configure Escalation Deployment](/NVIDIA-NeMo/Switchyard/how-can-one-reproduce-the-terminal-bench-2-1-benchmark-numbers-for-switchyard-and-what-configurations-are-included-in-the-escalation-deployment-toml)

Reproduce Terminal-Bench 2.1 benchmark results for Switchyard. Learn how to prepare the dataset, establish a baseline, and configure escalation deployment using the provided TOML file.

- Tags: how-to-guide
- Published: 2026-09-13

### [Purpose of fallback_base_url in Switchyard: Handling Unmatched HTTP Requests](/NVIDIA-NeMo/Switchyard/what-is-the-purpose-of-the-fallback-base-url-in-switchyard-and-how-are-http-requests-that-do-not-match-any-routes-handled)

Learn the purpose of fallback_base_url in Switchyard for handling unmatched HTTP requests. Discover how this catch-all endpoint proxies requests when no route matches. Optimize your LLM routing.

- Tags: how-to-guide
- Published: 2026-09-13

### [How a Passthrough Route in Switchyard Registers a Single Target Model ID Without Making a Routing Decision](/NVIDIA-NeMo/Switchyard/how-does-a-passthrough-route-in-switchyard-register-a-single-target-under-a-specific-model-id-without-making-a-routing-decision)

Learn how Switchyard's passthrough route registers a single target model ID, bypassing routing decisions. Optimize your NLU pipeline by forwarding requests directly.

- Tags: how-to-guide
- Published: 2026-09-13

### [Python Embedding Path vs Rust Algorithm::run_stream API in Switchyard: Key Differences](/NVIDIA-NeMo/Switchyard/what-is-the-difference-between-the-python-embedding-path-step-llmresponse-and-the-rust-algorithm-run-stream-api-in-switchyard)

Understand the key differences between Python embedding path (Step/LlmResponse) and Rust Algorithm::run_stream API in NVIDIA Switchyard. Learn how Python offers a type-checked facade over the Rust core.

- Tags: deep-dive
- Published: 2026-09-13

### [How Switchyard Routes `reasoning_effort` and `extra_body` Parameters to the Network](/NVIDIA-NeMo/Switchyard/how-do-reasoning-effort-and-extra-body-parameters-on-a-switchyard-target-reach-the-network-and-under-what-circumstances-are-they-dropped)

Discover how Switchyard routes reasoning_effort and extra_body parameters to the network. Learn when these parameters are dropped due to capability validation or existing request keys.

- Tags: internals
- Published: 2026-09-13

### [Switchyard Runner Duplicate Model ID Warning: Causes and Solutions](/NVIDIA-NeMo/Switchyard/why-does-the-switchyard-runner-issue-a-warning-for-duplicate-model-ids-on-the-same-llm-client-and-how-can-this-be-resolved)

Resolve Switchyard runner duplicate model ID warnings. Learn why this occurs on the same llm_client and how to fix it for deterministic routing.

- Tags: troubleshooting
- Published: 2026-09-13

### [Prometheus Counters Emitted by Switchyard-Server: Routing Overhead vs Latency Breakdown](/NVIDIA-NeMo/Switchyard/what-prometheus-counters-are-emitted-by-switchyard-server-and-how-do-they-differentiate-between-routing-overhead-and-latency)

Explore Switchyard-server's Prometheus counters. Understand how switchyard.routing_overhead_ms and switchyard.model_call_latency_ms differentiate routing overhead from model latency for optimized performance.

- Tags: performance
- Published: 2026-09-13

### [HTTP Endpoints Exposed by the Standalone Switchyard Proxy: Complete API Reference](/NVIDIA-NeMo/Switchyard/what-are-the-http-endpoints-exposed-by-the-standalone-switchyard-proxy-and-what-is-the-function-of-each-endpoint)

Explore the standalone Switchyard proxy's HTTP endpoints including OpenAI compatible chat completions, Anthropic APIs, and Prometheus metrics. Discover each endpoint's function in our API reference.

- Tags: api-reference
- Published: 2026-09-13

### [How Switchyard’s LiteLLM Integration Connects Routing to LiteLLM’s Router/Proxy](/NVIDIA-NeMo/Switchyard/how-does-the-litellm-integration-in-switchyards-examples-directory-connect-switchyard-routing-to-litellms-router-proxy)

Discover how Switchyard's LiteLLM integration connects routing to LiteLLM's Router/proxy. Learn how plugins select optimal models and rewrite requests for seamless AI routing.

- Tags: how-to-guide
- Published: 2026-09-13

### [How to Build, Package, and Configure the NVIDIA NeMo Relay Plugin for Switchyard to Use a Specific routes.toml Deployment](/NVIDIA-NeMo/Switchyard/how-is-the-nemo-relay-plugin-for-switchyard-built-packaged-and-configured-to-point-to-a-specific-routes-toml-deployment)

Learn to build, package, and configure the NVIDIA NeMo Relay plugin for Switchyard. Point your routes.toml deployment and dynamically route LLM requests with this Rust cdylib.

- Tags: how-to-guide
- Published: 2026-09-13

### [How Switchyard's Protocol Translation Layer Converts Between OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages](/NVIDIA-NeMo/Switchyard/how-does-switchyards-protocol-translation-layer-handle-conversions-between-openai-chat-completions-openai-responses-and-anthropic-messages)

Discover how Switchyard's protocol translation layer converts OpenAI Chat Completions, Responses, and Anthropic Messages to a vendor-agnostic format using codecs and re-encodes them for target providers.

- Tags: deep-dive
- Published: 2026-09-13

### [Understanding the target_selector Policy in NVIDIA Switchyard for Custom Multi‑Target Routing](/NVIDIA-NeMo/Switchyard/what-is-the-target-selector-policy-in-switchyard-and-how-can-it-be-utilized-for-custom-multi-target-routing-with-multiple-models)

Explore the target_selector policy in NVIDIA Switchyard to implement custom multi-target routing. Map JSON pointer categories to runtime models for deterministic, chained model selection.

- Tags: deep-dive
- Published: 2026-09-13

### [How Switchyard Implements Sub-Agent-Aware Routing with Passthrough and StageRouter Components](/NVIDIA-NeMo/Switchyard/how-is-sub-agent-aware-routing-implemented-in-switchyard-using-passthrough-and-stage-router-components)

Discover how Switchyard enables sub-agent-aware routing with passthrough and StageRouter. Learn how to delegate child work to isolated model groups for efficient processing.

- Tags: internals
- Published: 2026-09-13

### [Switchyard Advisor Gate Approval Loop: Mechanism and Signals](/NVIDIA-NeMo/Switchyard/explain-the-approval-loop-mechanism-of-the-advisor-gate-in-switchyard-and-the-signals-used-to-send-the-executor-back)

Discover the advisor gate approval loop in NVIDIA NeMo Switchyard. Learn how it uses advisor models and signals to approve or redo executor turns, ensuring robust workflow management.

- Tags: deep-dive
- Published: 2026-09-13

### [Switchyard Escalation Router: Trigger Conditions and Mode Differences Explained](/NVIDIA-NeMo/Switchyard/what-conditions-trigger-escalation-in-switchyards-escalation-router-and-how-does-mode-escalation-differ-from-a-standard-llm-classifier)

Understand Switchyard escalation router trigger conditions and mode differences. Learn how escalation differs from standard LLM classifiers for improved conversational AI.

- Tags: deep-dive
- Published: 2026-09-13

### [How the Switchyard Stage Router Chooses Between Tool-Response Pattern Matching and an LLM Judge](/NVIDIA-NeMo/Switchyard/how-does-the-stage-router-in-switchyard-make-decisions-between-tool-response-pattern-matching-and-using-an-llm-judge)

Discover how the Switchyard stage router selects between tool-response pattern matching and an LLM judge. Learn about its decision-making process for optimal routing.

- Tags: how-to-guide
- Published: 2026-09-13

### [How Switchyard Validates Version-1 TOML Deployment Configurations: Schema, LLM Clients, and Routing Rules](/NVIDIA-NeMo/Switchyard/how-does-the-version-1-toml-deployment-schema-in-switchyard-including-llm-clients-targets-and-routes-validate-a-switchyard-configuration)

Learn how Switchyard validates version-1 TOML deployment configurations. Discover schema checks, LLM client validation, and routing rule consistency for robust deployments.

- Tags: deep-dive
- Published: 2026-09-13

### [RuntimeModels and Scope Separation in NVIDIA NeMo Switchyard: Isolating Parent and Sub-Agent Traffic](/NVIDIA-NeMo/Switchyard/what-is-runtimemodels-in-switchyard-and-how-does-scope-separation-maintain-isolation-between-parent-and-sub-agent-traffic)

Discover how Switchyard's RuntimeModels and Scope separation isolate parent and sub-agent traffic. Learn about per-request dictionaries and independent routing contexts for robust agent communication.

- Tags: deep-dive
- Published: 2026-09-13

### [Switchyard RoutingOutcome Structure: How Model Selection Drives the Answer Call](/NVIDIA-NeMo/Switchyard/how-is-the-routing-outcome-structure-defined-in-switchyard-and-what-are-the-implications-of-selected-model-ids-request-response-and-metadata-for-the-answer-call)

Understand Switchyard's RoutingOutcome structure. Learn how selected model IDs, request, response, and metadata influence the answer call for efficient routing.

- Tags: deep-dive
- Published: 2026-09-13

### [Understanding the Step Protocol in Switchyard: CallModel, Done, and the respond Contract](/NVIDIA-NeMo/Switchyard/what-is-the-step-protocol-in-switchyard-including-callmodel-and-done-and-what-is-the-contract-for-callmodel-respond)

Explore the Switchyard Step protocol including CallModel and Done. Understand the CallModel respond contract for seamless routing and fallback.

- Tags: deep-dive
- Published: 2026-09-13

### [How to Integrate NVIDIA NeMo Switchyard Algorithm::run_stream with a Custom Harness](/NVIDIA-NeMo/Switchyard/how-does-nvidia-nemo-switchyard-algorithm-run-stream-function-operate-and-how-can-it-be-integrated-with-a-custom-harness)

Learn how NVIDIA NeMo Switchyard's Algorithm::run_stream works and integrates with custom harnesses. Interleave routing logic with external LLM execution effectively.

- Tags: how-to-guide
- Published: 2026-09-13

### [How Switchyard Handles Context Window Overflow: Automatic Fallback to Capable Models](/NVIDIA-NeMo/Switchyard/how-switchyard-manage-context-window-overflow-capable-model-requires-more-tokens-efficient-model)

Discover how Switchyard gracefully handles context window overflow, automatically falling back to capable models when efficient models exceed token limits, ensuring uninterrupted performance.

- Tags: deep-dive
- Published: 2026-09-12

### [How Switchyard's Soak-Test Methodology Evaluates Routing Overhead and Latency Under Load](/NVIDIA-NeMo/Switchyard/how-switchyard-soak-test-methodology-evaluate-routing-overhead-latency-load)

Switchyard's soak-test measures routing overhead and latency by comparing routed deployments to direct-backend baselines using deterministic traffic scenarios and tools like OHA and AIPerf.

- Tags: performance
- Published: 2026-09-12

### [How to Configure Prometheus Metrics and Enable the /metrics and /v1/stats Endpoints in Switchyard](/NVIDIA-NeMo/Switchyard/how-configure-prometheus-metrics-enable-metrics-v1-stats-endpoints-switchyard)

Configure Switchyard for Prometheus metrics and enable /metrics and /v1/stats endpoints. Simply set metrics and stats to true in your TOML file and restart.

- Tags: how-to-guide
- Published: 2026-09-12

### [How Switchyard Resolves Route IDs to Targets for OpenAI-Compatible Endpoints](/NVIDIA-NeMo/Switchyard/how-switchyard-server-resolve-route-ids-targets-v1-chat-completions-v1-messages-v1-responses)

Learn how Switchyard resolves route IDs to targets for OpenAI-compatible endpoints like v1 chat completions. Discover the routing algorithm and target selection process.

- Tags: deep-dive
- Published: 2026-09-12

### [How to Develop a LiteLLM Routing Plugin for Switchyard Using switchyard_litellm](/NVIDIA-NeMo/Switchyard/how-develop-litellm-routing-plugin-switchyard-using-switchyard-litellm)

Learn to develop a LiteLLM routing plugin for Switchyard. Implement a CustomLogger, intercept requests with _request, and use build_request_patch to modify them with Switchyard's Python bindings.

- Tags: how-to-guide
- Published: 2026-09-12

### [How to Build a Custom NeMo Relay Plugin Bundle for Switchyard Using relay-plugin.toml and a Digest](/NVIDIA-NeMo/Switchyard/how-build-custom-nemo-relay-plugin-bundle-switchyard-relay-plugin-toml-digest)

Learn to build a custom NeMo Relay plugin bundle for Switchyard. Compile the crate, calculate a digest, and package it with relay-plugin.toml using package_bundle.py.

- Tags: how-to-guide
- Published: 2026-09-12

### [How to Configure OpenTelemetry Spans in Switchyard to Trace Routing Decisions per Turn](/NVIDIA-NeMo/Switchyard/how-configure-opentelemetry-spans-switchyard-tracing-routing-decisions-per-turn)

Configure OpenTelemetry spans in Switchyard to trace routing decisions per turn. Learn how Switchyard instruments each turn and captures routing details in spans.

- Tags: how-to-guide
- Published: 2026-09-12

### [How to Add a Custom Processor to Switchyard for Observing or Modifying Request State During Streaming](/NVIDIA-NeMo/Switchyard/how-add-custom-processor-switchyard-observing-modifying-request-state-streaming)

Learn to add a custom Processor to Switchyard to observe or modify request state during streaming. Implement the Processor trait and use with_processor for efficient stream manipulation.

- Tags: how-to-guide
- Published: 2026-09-12

### [How to Implement a Custom Algorithm in Switchyard by Extending the Algorithm and Classifier Traits](/NVIDIA-NeMo/Switchyard/how-implement-custom-algorithm-switchyard-extending-algorithm-classifier-traits)

Implement custom algorithms in NVIDIA Switchyard by extending Algorithm and Classifier traits. Learn to define name() and route() methods and register your implementation for advanced routing.

- Tags: how-to-guide
- Published: 2026-09-12

### [How to Migrate from LlmTarget API to Step/LlmResponse API in Switchyard](/NVIDIA-NeMo/Switchyard/how-migrate-older-llmtarget-api-newer-step-llmresponse-api-switchyard)

Migrate from LlmTarget API to Switchyard's Step/LlmResponse API. Update imports, use an async loop for run_stream, and wrap outputs for aggregated or streamed responses. Learn the upgrade process.

- Tags: migration-guide
- Published: 2026-09-12

### [#LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard: Understanding Streaming Behavior](/NVIDIA-NeMo/Switchyard/difference-streaming-behavior-llmresponsestream-aggllmresponse-switchyard-algorithms)

Understand LlmResponseStream vs AggLlmResponse in NVIDIA Switchyard. Discover how LlmResponseStream offers real-time tokens for low latency while AggLlmResponse provides complete output after generation.

- Tags: deep-dive
- Published: 2026-09-12

### [How Switchyard's IR Translation Layer Normalizes OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages](/NVIDIA-NeMo/Switchyard/how-switchyard-ir-translation-layer-normalize-openai-anthropic-messages)

Learn how Switchyard's IR translation layer normalizes OpenAI Chat Completions and Anthropic Messages into a unified format, enabling seamless integration and efficient processing of LLM outputs.

- Tags: internals
- Published: 2026-09-12

### [Step::CallModel and Step::Done in Switchyard's Host-Side Orchestration Loop](/NVIDIA-NeMo/Switchyard/role-step-callmodel-step-done-switchyard-host-side-orchestration-loop)

Understand Step::CallModel and Step::Done in Switchyard's host orchestration loop. Learn how these control messages decouple routing from I/O and manage LLM invocation and request termination.

- Tags: internals
- Published: 2026-09-12

### [How Switchyard's Driver::call_model Offloads Model Execution to the Host Harness](/NVIDIA-NeMo/Switchyard/how-driver-call-model-switchyard-offload-model-execution-host-harness)

Learn how Switchyard's Driver::call_model offloads LLM inference to a host harness. Discover how it publishes steps via promise-based channels for efficient, I/O-free routing.

- Tags: internals
- Published: 2026-09-12

### [How PickerMode Options Interact with Confidence Thresholds in Switchyard](/NVIDIA-NeMo/Switchyard/how-switchyard-pickermode-options-interact-confidence-thresholds)

Explore how Switchyard's PickerMode options like efficient_first, capable_first, and weighted interact with confidence thresholds to control routing decisions. Optimize your model's fallback tiers.

- Tags: deep-dive
- Published: 2026-09-12

### [How the target_selector Policy Routes Requests Among Multiple Targets in Switchyard](/NVIDIA-NeMo/Switchyard/how-switchyard-target-selector-policy-route-requests-multiple-targets)

Learn how the target_selector policy in Switchyard routes requests by mapping classifier verdicts to runtime models for deterministic classification. Understand its abstention mechanism.

- Tags: how-to-guide
- Published: 2026-09-12

### [How Switchyard's Sub-Agent-Aware Routing Isolates Traffic for Parent and Delegated Agents](/NVIDIA-NeMo/Switchyard/how-switchyard-sub-agent-aware-routing-isolate-traffic-parent-delegated-agents)

Discover how Switchyard's sub-agent-aware routing isolates traffic, keeping parent agent requests distinct from delegated sub-agent work for enhanced control and performance.

- Tags: deep-dive
- Published: 2026-09-12

### [Composite Routing in NVIDIA Switchyard: Combining LLM Classifier and Stage Router Cascades](/NVIDIA-NeMo/Switchyard/how-switchyard-implements-composite-routing-combining-llm-classifier-stage-router-cascades)

Discover how Switchyard integrates LLM classifiers with Stage Router cascades for advanced composite routing. Learn about the TierSetter processor and model pool determination.

- Tags: deep-dive
- Published: 2026-09-12

### [Switchyard Advisor Gate vs Stage Routing: Functional Differences in Plan Evaluation](/NVIDIA-NeMo/Switchyard/functional-difference-switchyard-advisor-gate-stage-routing-plan-evaluation)

Understand the functional differences between Switchyard Advisor Gate and Stage Routing. Advisor Gate offers quality control checkpoints, while Stage Routing dynamically selects models for plan evaluation.

- Tags: deep-dive
- Published: 2026-09-12

### [How Switchyard Escalation Routing Detects and Handles Model Response Failures](/NVIDIA-NeMo/Switchyard/how-switchyard-escalation-routing-detect-handle-model-response-failures)

Learn how Switchyard escalation routing detects and handles model response failures by categorizing errors and triggering fail open mechanisms to maintain session streaks and preserve weak responses.

- Tags: how-to-guide
- Published: 2026-09-12

### [How Switchyard's StageRouter Chooses Between Efficient and Capable Tiers Using Tool Signals](/NVIDIA-NeMo/Switchyard/switchyard-stagerouter-mechanism-choose-efficient-capable-tiers-tool-signals)

Switchyard's StageRouter intelligently selects between efficient and capable tiers using tool signals and picker modes. Learn how it optimizes performance and reliability.

- Tags: deep-dive
- Published: 2026-09-12

### [How Switchyard RuntimeModels Differentiates Parent and Sub-Agent Model Groups](/NVIDIA-NeMo/Switchyard/how-does-switchyard-runtime-models-differentiate-parent-sub-agent-model-groups)

Learn how Switchyard's RuntimeModels separates parent and sub-agent model groups using distinct HashMaps for isolated management and access. Understand the differentiation for robust agent architecture.

- Tags: internals
- Published: 2026-09-12

### [How NVIDIA Switchyard’s Algorithm::run_stream Handles Internal Routing Decisions](/NVIDIA-NeMo/Switchyard/how-does-nvidia-switchyard-algorithm-run-stream-handle-internal-routing-decisions)

Learn how NVIDIA Switchyard's Algorithm::run_stream manages internal routing decisions by spawning Drivers for efficient, non-blocking I/O and model execution.

- Tags: internals
- Published: 2026-09-12

### [How to Uninstall Switchyard: Complete Removal Guide for Python and Rust Components](/NVIDIA-NeMo/Switchyard/how-to-uninstall-switchyard)

Uninstall Switchyard completely from your system. Learn how to remove Python and Rust components with clear, step-by-step instructions for a full cleanup.

- Tags: how-to-guide
- Published: 2026-09-11

### [Switchyard Development Best Practices: A Guide to Modular LLM Routing](/NVIDIA-NeMo/Switchyard/what-are-the-best-practices-for-switchyard-development)

Discover Switchyard development best practices for modular LLM routing. Learn about async algorithms, type safety, and testing for reliable Rust and Python integrations.

- Tags: best-practices
- Published: 2026-09-11

### [How to Manage Configurations in Switchyard: A Complete TOML-Based Guide](/NVIDIA-NeMo/Switchyard/how-to-manage-configurations-in-switchyard)

Master Switchyard configuration management with this complete TOML guide. Learn how Switchyard uses TOML to create typed DeploymentConfig objects for inference and HTTP server.

- Tags: how-to-guide
- Published: 2026-09-11

### [Typical Workflow for Using Switchyard: Three Integration Methods Explained](/NVIDIA-NeMo/Switchyard/what-is-the-typical-workflow-for-using-switchyard)

Discover the typical Switchyard workflow integrating LLM pipelines through NeMo Relay, Rust/Python libraries, or a proxy server. Learn common routing algorithms for seamless integration.

- Tags: how-to-guide
- Published: 2026-09-11

### [How to Get Support for Switchyard: Official Channels and Best Practices](/NVIDIA-NeMo/Switchyard/how-to-get-support-for-switchyard)

Get Switchyard support via GitHub Issues for bugs and features, consult documentation for troubleshooting, or use the NVIDIA PSIRT channel for security vulnerabilities.

- Tags: support
- Published: 2026-09-11

### [Switchyard Dependencies: Complete Guide to the NVIDIA LLM Router Crate Ecosystem](/NVIDIA-NeMo/Switchyard/what-are-the-dependencies-for-switchyard)

Explore the comprehensive dependencies of the NVIDIA LLM Router crate within Switchyard. Understand its core async libraries and internal workspace crates for efficient LLM management.

- Tags: deep-dive
- Published: 2026-09-11

### [How to Debug Issues in Switchyard: OpenTelemetry Metrics and Tracing Guide](/NVIDIA-NeMo/Switchyard/how-to-debug-issues-in-switchyard)

Debug Switchyard issues effectively. Utilize OpenTelemetry metrics and tracing to inspect routing and model selections in real-time. Enable verbose logging for deeper insights.

- Tags: how-to-guide
- Published: 2026-09-11

### [Are There Pre-Trained Models Available for Switchyard? A Technical Guide to LLM Routing](/NVIDIA-NeMo/Switchyard/are-there-pre-trained-models-available-for-switchyard)

Discover if Switchyard offers pre-trained models. Learn how this LLM routing layer connects to external services like OpenAI and NVIDIA NeMo Relay for efficient inference.

- Tags: how-to-guide
- Published: 2026-09-11

### [How to Deploy Switchyard Models: Server and Library Configuration Guide](/NVIDIA-NeMo/Switchyard/how-to-deploy-switchyard-models)

Learn to deploy Switchyard models efficiently. Configure clients, targets, and routes in TOML, then run the server or embed the library in your Rust app.

- Tags: how-to-guide
- Published: 2026-09-11

### [Switchyard Performance Optimizations: 8 Ways to Speed Up LLM Routing](/NVIDIA-NeMo/Switchyard/what-performance-optimizations-are-available-in-switchyard)

Discover 8 Switchyard performance optimizations to speed up LLM routing. Reduce latency and maximize throughput with intelligent routing, token reuse, and more.

- Tags: performance
- Published: 2026-09-11

### [How to Integrate Switchyard with NVIDIA AI Tools: 3 Integration Methods Explained](/NVIDIA-NeMo/Switchyard/how-to-integrate-switchyard-with-other-nvidia-ai-tools)

Discover how to integrate Switchyard with NVIDIA AI tools using three methods: NeMo Relay plugin, direct library embedding, or an OpenAI-compatible proxy. Optimize your AI workflows.

- Tags: how-to-guide
- Published: 2026-09-11

### [Does Switchyard Support Distributed Training? A Deep Dive into the NVIDIA-NeMo Routing Engine](/NVIDIA-NeMo/Switchyard/does-switchyard-support-distributed-training)

Discover if Switchyard supports distributed training. Learn how this NVIDIA-NeMo routing engine focuses solely on LLM inference, not training or data parallelism.

- Tags: deep-dive
- Published: 2026-09-11

### [How to Fine-Tune Models with Switchyard: A Complete Integration Guide](/NVIDIA-NeMo/Switchyard/how-to-fine-tune-models-with-switchyard)

Learn how to fine-tune models with Switchyard. Integrate custom models by deploying checkpoints and routing traffic with this comprehensive guide. Get started today!

- Tags: how-to-guide
- Published: 2026-09-11

### [Switchyard Components: The Complete Architecture of NVIDIA’s LLM Router](/NVIDIA-NeMo/Switchyard/what-are-the-main-components-of-switchyard)

Explore the core components of Switchyard a Rust-based LLM router. Understand its architecture including libsy llm-client runner server Python bindings protocol and translation for efficient provider-agnostic routing.

- Tags: architecture
- Published: 2026-09-11

### [How to Use Switchyard for Custom Model Development: A Complete Guide](/NVIDIA-NeMo/Switchyard/how-to-use-switchyard-for-custom-model-development)

Learn custom model development with Switchyard. This guide shows how Switchyard routes LLM requests to targets without client code changes, enabling flexible and efficient development.

- Tags: how-to-guide
- Published: 2026-09-11

### [Where to Find Switchyard Documentation: Complete Guide to NVIDIA-NeMo's LLM Router](/NVIDIA-NeMo/Switchyard/where-can-i-find-documentation-for-switchyard)

Find NVIDIA NeMo Switchyard documentation easily. Explore the official repository for installation guides and configuration schemas to get started with your LLM router.

- Tags: getting-started
- Published: 2026-09-11

### [How to Run the Examples in the Switchyard Repository](/NVIDIA-NeMo/Switchyard/how-to-run-examples-in-switchyard)

Run Switchyard examples easily. Execute Python LibSy driver for tests or deploy LiteLLM routing plugin with Docker Compose for full proxy integration. Get started now.

- Tags: getting-started
- Published: 2026-09-11

### [Switchyard Repository Licensing: Apache 2.0 Terms and Compliance Guide](/NVIDIA-NeMo/Switchyard/what-is-the-licensing-for-switchyard-repository)

Understand the NVIDIA-NeMo Switchyard repository licensing under Apache 2.0. Learn terms for commercial use modification and distribution.

- Tags: licensing-guide
- Published: 2026-09-11

### [What Programming Languages Does Switchyard Support? Native Rust and Python Bindings Explained](/NVIDIA-NeMo/Switchyard/what-programming-languages-does-switchyard-support)

Discover Switchyard's programming language support. Explore native Rust and Python bindings for seamless integration or use the standalone server with any language.

- Tags: getting-started
- Published: 2026-09-11

### [How to Set Up a Development Environment for Switchyard](/NVIDIA-NeMo/Switchyard/how-to-set-up-development-environment-for-switchyard)

Easily set up a Switchyard development environment. Install Python, Rust, and uv, then sync dependencies to get started fast. Build with Switchyard today.

- Tags: getting-started
- Published: 2026-09-11

### [System Requirements for NVIDIA NeMo Switchyard: Hardware, Software, and Build Prerequisites](/NVIDIA-NeMo/Switchyard/what-are-the-system-requirements-for-switchyard)

Discover NVIDIA NeMo Switchyard system requirements. Learn about essential hardware like x86_64-v3 CPUs, and software including Python 3.10+ and Rust 1.96.1+ for optimal performance.

- Tags: system-requirements
- Published: 2026-09-11

### [How to Install Switchyard from NVIDIA-NeMo: Python, Server, and Rust Crate Setup](/NVIDIA-NeMo/Switchyard/how-to-install-switchyard-from-nvidia-nemo)

Install Switchyard from NVIDIA NeMo easily. Follow guides for Python pip, Rust cargo, and server binary setups to integrate Switchyard into your projects today.

- Tags: getting-started
- Published: 2026-09-11

### [Switchyard Library APIs: Server Management and LLM Routing Reference](/NVIDIA-NeMo/Switchyard/what-are-the-key-apis-in-the-switchyard-library)

Explore NVIDIA Switchyard library APIs for server management and LLM routing. Discover key Python packages switchyard_rust.server and switchyard.libsy for building routing algorithms.

- Tags: api-reference
- Published: 2026-08-23

### [Understanding the Switchyard Architecture: NVIDIA's LLM Traffic Proxy Explained](/NVIDIA-NeMo/Switchyard/how-to-understand-the-switchyard-architecture)

Understand Switchyard, NVIDIA's LLM traffic proxy. Normalize API requests, route them with algorithms, and translate responses seamlessly for provider-agnostic LLM interactions.

- Tags: architecture
- Published: 2026-08-23

### [Best Practices for Developing with Switchyard: A Comprehensive Guide](/NVIDIA-NeMo/Switchyard/what-are-the-best-practices-for-developing-with-switchyard)

Master Switchyard development with our comprehensive guide. Learn best practices for dependency management testing strict typing Rust safety and commit conventions to build robust applications.

- Tags: best-practices
- Published: 2026-08-23

### [How to Set Up CI/CD for Switchyard Projects: A Complete GitHub Actions Guide](/NVIDIA-NeMo/Switchyard/how-to-set-up-ci-cd-for-switchyard-projects)

Master CI/CD for Switchyard projects using GitHub Actions. Automate Python linting, Rust builds, and Maturin wheel packaging for robust validation on every push.

- Tags: how-to-guide
- Published: 2026-08-23

### [Switchyard Performance Considerations: Latency, Throughput, and Scaling Explained](/NVIDIA-NeMo/Switchyard/what-are-the-performance-considerations-for-switchyard)

Understand Switchyard performance considerations: latency, throughput, and scaling. Discover minimal overhead and linear scaling for LLM requests.

- Tags: performance
- Published: 2026-08-23

### [Switchyard Release Cycle: How Tag-Driven Automation Publishes to PyPI and crates.io](/NVIDIA-NeMo/Switchyard/what-is-the-release-cycle-for-switchyard)

Discover the Switchyard release cycle. Learn how tag-driven automation publishes to PyPI and crates.io, ensuring efficient and reliable updates for Python and Rust.

- Tags: internals
- Published: 2026-08-23

### [How to Report a Bug in Switchyard: The Complete Contributor Guide](/NVIDIA-NeMo/Switchyard/how-to-report-a-bug-in-switchyard)

Learn how to report a bug in Switchyard effectively. Follow our guide to create a GitHub Issue with a reproducible example and environment details for quick resolution.

- Tags: how-to-guide
- Published: 2026-08-23

### [What License Governs NVIDIA Switchyard? Apache 2.0 Explained](/NVIDIA-NeMo/Switchyard/what-is-the-licensing-for-nvidia-switchyard)

Discover the licensing for NVIDIA Switchyard. It uses the permissive Apache License 2.0, enabling free use, modification, and distribution. Learn the key terms.

- Tags: licensing-information
- Published: 2026-08-23

### [How to Debug Switchyard Applications: End-to-End Logging for NVIDIA NeMo’s Python-Rust Stack](/NVIDIA-NeMo/Switchyard/how-to-debug-switchyard-applications)

Debug Switchyard applications effectively with end-to-end logging for NVIDIA NeMo's Python-Rust stack. Learn to trace requests, mock tests with respx, and validate Rust code.

- Tags: how-to-guide
- Published: 2026-08-23

### [How to Build Switchyard from Source: Complete Installation Guide](/NVIDIA-NeMo/Switchyard/how-to-build-switchyard-from-source)

Build Switchyard from source with our complete guide. Install Rust and uv, clone the repo, and compile the Rust workspace for a seamless setup.

- Tags: how-to-guide
- Published: 2026-08-23

### [Switchyard Development Dependencies: A Complete Guide to the Rust and Python Toolchain](/NVIDIA-NeMo/Switchyard/what-are-the-dependencies-for-switchyard-development)

Master Switchyard development dependencies with our guide to the Rust and Python toolchain. Learn about Rust crates like tokio, axum, serde, pyo3, and Python's pytest for efficient development.

- Tags: getting-started
- Published: 2026-08-23

### [How to Run Unit Tests in Switchyard: A Complete Guide](/NVIDIA-NeMo/Switchyard/how-to-run-unit-tests-in-switchyard)

Learn to run unit tests in Switchyard with this guide. Install uv, sync dependencies, and execute pytest tests easily without API keys or external services.

- Tags: how-to-guide
- Published: 2026-08-23

### [How to Contribute to NVIDIA Switchyard: A Complete Step-by-Step Guide](/NVIDIA-NeMo/Switchyard/how-to-contribute-to-nvidia-switchyard)

Learn how to contribute to NVIDIA Switchyard with this step-by-step guide. Fork the repo, create a branch, test your changes, and submit a pull request to enhance this powerful tool.

- Tags: how-to-guide
- Published: 2026-08-23

### [Main Directories in the NVIDIA-NeMo Switchyard Repository: A Complete Guide](/NVIDIA-NeMo/Switchyard/what-are-the-main-directories-in-the-switchyard-repository)

Explore the main directories of the NVIDIA-NeMo Switchyard repository. Understand the Python API, Rust core, tests, examples, and more for efficient LLM orchestration.

- Tags: how-to-guide
- Published: 2026-08-23

### [How Is the Switchyard Codebase Structured? Inside NVIDIA-NeMo's Python-Rust Hybrid Architecture](/NVIDIA-NeMo/Switchyard/how-is-the-switchyard-codebase-structured)

Explore the NVIDIA NeMo Switchyard codebase structure, a Python Rust hybrid architecture. Discover how it separates API ergonomics from high-performance logic for efficient routing and translation.

- Tags: internals
- Published: 2026-08-23

### [Programming Languages Used in Switchyard: Rust Core, Python API, and More](/NVIDIA-NeMo/Switchyard/what-programming-languages-are-used-in-switchyard)

Discover the programming languages powering NVIDIA Switchyard. Learn how Rust drives the core engine and Python powers the API, optimizing your projects.

- Tags: deep-dive
- Published: 2026-08-23

### [What Are the Core Components of NVIDIA Switchyard? A Complete Architecture Breakdown](/NVIDIA-NeMo/Switchyard/what-are-the-core-components-of-nvidia-switchyard)

Explore the core components of NVIDIA Switchyard. Understand the architecture of this Rust and Python LLM routing layer including server, libsy, protocol, translation, and llm client.

- Tags: architecture
- Published: 2026-08-23

### [How to Install NVIDIA Switchyard: Python Package, Rust Server, and Library Integration](/NVIDIA-NeMo/Switchyard/how-to-install-nvidia-switchyard)

Install NVIDIA Switchyard easily via pip for Python, cargo for the Rust server, or integrate individual crates into your Rust projects. Get started with Switchyard today.

- Tags: getting-started
- Published: 2026-08-23

### [What Is NVIDIA Switchyard? A Rust-Based LLM Proxy and Router](/NVIDIA-NeMo/Switchyard/what-is-nvidia-switchyard)

Discover NVIDIA Switchyard, a Rust proxy that translates OpenAI and Anthropic API formats. Route requests to multiple LLM backends with pluggable algorithms.

- Tags: getting-started
- Published: 2026-08-23

### [Switchyard Route Types and Their Basic Configurations](/NVIDIA-NeMo/Switchyard/switchyard-supported-route-types-configuration)

Discover Switchyard route types like llm_classifier, stage_router, and more. Learn their basic configurations to optimize your AI workflows.

- Tags: how-to-guide
- Published: 2026-08-22

### [How to Deploy and Use a Self-Hosted OpenAI-Compatible Model Server (vLLM) with Switchyard](/NVIDIA-NeMo/Switchyard/deploy-self-hosted-model-with-switchyard-vllm)

Easily deploy and use a self-hosted OpenAI-compatible model server like vLLM with Switchyard. This routing layer requires no code changes for seamless integration of local models via simple TOML configuration.

- Tags: how-to-guide
- Published: 2026-08-22

### [How Switchyard Handles Bidirectional Translation for Streaming and Buffered LLM Responses](/NVIDIA-NeMo/Switchyard/switchyard-bidirectional-response-translation)

Discover how Switchyard masterfully translates LLM responses. It decodes any wire format to a neutral IR, then re-encodes for streaming or buffered output, ensuring seamless bidirectional translation.

- Tags: internals
- Published: 2026-08-22

### [Supported API Formats for LLM Clients in Switchyard and Their Upstream Endpoints](/NVIDIA-NeMo/Switchyard/switchyard-supported-api-formats-upstream-endpoints)

Discover Switchyard's supported API formats for LLM clients including OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages. Learn about their upstream endpoints.

- Tags: api-reference
- Published: 2026-08-22

### [How to Define a Target Model Endpoint in Switchyard’s TOML Configuration](/NVIDIA-NeMo/Switchyard/define-target-model-endpoint-switchyard-toml)

Learn to define a target model endpoint in Switchyard using TOML configuration. Set up your upstream model identifier and client connection for seamless routing.

- Tags: how-to-guide
- Published: 2026-08-22

### [How to Configure an LLM Client in Switchyard for OpenRouter](/NVIDIA-NeMo/Switchyard/configure-llm-client-switchyard-openrouter)

Learn to configure an LLM client in Switchyard for OpenRouter. Route LLM requests via direct Python or TOML server setup for seamless integration.

- Tags: how-to-guide
- Published: 2026-08-22

### [TOML Configuration File for Switchyard Deployments: Complete Schema Guide](/NVIDIA-NeMo/Switchyard/switchyard-toml-configuration-structure)

Understand the TOML configuration file structure for Switchyard deployments. Learn about LLM clients, model targets, and routing endpoints in this comprehensive schema guide.

- Tags: api-reference
- Published: 2026-08-22

### [How to Configure Sub-Agent Routing in Switchyard: A Complete Guide](/NVIDIA-NeMo/Switchyard/configure-sub-agent-routing-switchyard)

Learn to configure sub-agent routing in Switchyard using the subagents table in your TOML. Route child tasks to specialized classifiers without altering parent-agent policies.

- Tags: how-to-guide
- Published: 2026-08-22

### [How Advisor Gate Routing Works in Switchyard: Executor-Reviewer Architecture Explained](/NVIDIA-NeMo/Switchyard/advisor-gate-routing-advisor-reviewer)

Discover how Advisor Gate Routing in Switchyard uses an executor and advisor LLM to ensure response quality through automated approval gates.

- Tags: deep-dive
- Published: 2026-08-22

### [Escalation Router Routing in NVIDIA-NeMo Switchyard: Automatic Tier Escalation Strategy](/NVIDIA-NeMo/Switchyard/escalation-router-routing-strategy-switchyard)

Explore Escalation Router Routing in NVIDIA-NeMo Switchyard. Automatically elevate conversations to strong models using LLM judges for efficient cost management.

- Tags: deep-dive
- Published: 2026-08-22

### [How the LLM Classifier Routing Escalation Mode Works in Switchyard](/NVIDIA-NeMo/Switchyard/llm-classifier-escalation-mode-explained)

Understand Switchyard's LLM Classifier Routing escalation mode. Learn how it uses an efficient LLM, judge LLM, and configurable criteria for intelligent response routing.

- Tags: how-to-guide
- Published: 2026-08-22

### [LLM Classifier Routing in Switchyard: 3 Modes Explained](/NVIDIA-NeMo/Switchyard/llm-classifier-routing-modes-switchyard)

Explore the three LLM classifier routing modes in Switchyard: Capability, Escalation, and Custom. Learn how to direct requests between models using LlmClassifierConfig.

- Tags: deep-dive
- Published: 2026-08-22

### [How to Configure Weighted Random Routing in Switchyard for A/B Testing](/NVIDIA-NeMo/Switchyard/configure-weighted-random-routing-switchyard)

Configure weighted random routing in Switchyard for A/B testing. Distribute traffic across models using TOML or Python API for efficient experimentation.

- Tags: how-to-guide
- Published: 2026-08-22

### [Switchyard Request Lifecycle: From TCP Ingress to LLM Response](/NVIDIA-NeMo/Switchyard/switchyard-request-lifecycle-steps)

Explore the Switchyard request lifecycle. See how it handles LLM requests from TCP ingress to Axum pipeline routing, client forwarding, and response serialization.

- Tags: deep-dive
- Published: 2026-08-22

### [How Python Applications Interact with Switchyard's Routing Algorithms Using switchyard-py](/NVIDIA-NeMo/Switchyard/python-integration-with-switchyard-py)

Learn how Python applications interact with Switchyard's routing algorithms using the switchyard-py package. Explore PyO3 bindings for seamless integration and efficient routing.

- Tags: how-to-guide
- Published: 2026-08-22

### [How the Switchyard Protocol Crate Defines Provider-Neutral Request and Response Types](/NVIDIA-NeMo/Switchyard/protocol-crate-provider-neutral-types)

Discover how the Switchyard protocol crate defines provider-neutral request and response types with pure-Rust structs for unified LLM interactions and lossless data handling.

- Tags: internals
- Published: 2026-08-22

### [How switchyard‑llm‑client Enables Upstream Model Calls in NVIDIA Switchyard](/NVIDIA-NeMo/Switchyard/switchyard-llm-client-upstream-calls)

Discover how switchyard-llm-client acts as an HTTP bridge to make authenticated, protocol-specific calls to upstream LLM models like OpenAI, Anthropic, and NVIDIA NIM.

- Tags: how-to-guide
- Published: 2026-08-22

### [How Switchyard-Translation Enables API Format Conversion for LLMs](/NVIDIA-NeMo/Switchyard/switchyard-translation-api-format-conversion)

Discover how Switchyard-translation, a Rust crate, converts API formats for LLMs. Seamlessly route requests between OpenAI, Anthropic, and other backends with bidirectional mapping.

- Tags: how-to-guide
- Published: 2026-08-22

### [Routing Algorithms in the libsy Crate: Complete Guide to LLM Traffic Orchestration in Switchyard](/NVIDIA-NeMo/Switchyard/libsy-routing-algorithms-list)

Explore nine routing algorithms in the libsy crate for LLM traffic orchestration. Discover Passthrough, Random, FallThrough, LlmTaskClassifier, and more for efficient request dispatch.

- Tags: deep-dive
- Published: 2026-08-22

### [How the switchyard-server Crate Functions as an HTTP Server for LLM Traffic](/NVIDIA-NeMo/Switchyard/switchyard-server-http-server-functionality)

Discover how the switchyard-server Rust crate acts as an HTTP server, efficiently proxying and routing LLM traffic with OpenAI/Anthropic API compatibility and Prometheus metrics.

- Tags: how-to-guide
- Published: 2026-08-22

### [Main Crates in the NVIDIA Switchyard Project: Complete Architecture Guide](/NVIDIA-NeMo/Switchyard/switchyard-crate-architecture-overview)

Explore the main Rust crates in the NVIDIA Switchyard project, an LLM traffic orchestration stack. Understand the architecture from the core engine to servers and Python bindings.

- Tags: architecture
- Published: 2026-08-22

### [How Switchyard Routes LLM Requests Across Different Providers](/NVIDIA-NeMo/Switchyard/how-does-switchyard-handle-llm-routing)

Discover how Switchyard routes LLM requests across providers using its Python-Rust architecture. Learn about configurable algorithms for intelligent target selection and delegation.

- Tags: how-to-guide
- Published: 2026-08-22

### [What Is NVIDIA Switchyard? A Rust‑Based Proxy for LLM Traffic Orchestration](/NVIDIA-NeMo/Switchyard/what-is-switchyard-and-its-purpose)

Explore NVIDIA Switchyard, a Rust proxy for LLM traffic. It simplifies protocol translation, intelligent routing, and observability for seamless backend integration.

- Tags: getting-started
- Published: 2026-08-22

### [How Switchyard Handles LLM Request Metadata: Session ID and Correlation ID](/NVIDIA-NeMo/Switchyard/how-switchyard-handles-llm-request-metadata)

Learn how Switchyard handles LLM request metadata, normalizing HTTP headers into a Metadata struct for session and correlation IDs, enabling distributed tracing and end-to-end observability.

- Tags: internals
- Published: 2026-08-21

### [Switchyard Requirements: Rust Toolchain, Python Version, and System Prerequisites](/NVIDIA-NeMo/Switchyard/what-are-requirements-running-switchyard)

Discover Switchyard requirements: Rust 1.96.1, Python 3.10+, and essential system tools like gcc and git are needed to get started. Optimize your setup now.

- Tags: getting-started
- Published: 2026-08-21

### [How to Monitor Per-Session Routing Statistics in Switchyard](/NVIDIA-NeMo/Switchyard/how-to-monitor-per-session-routing-statistics-switchyard)

Monitor per-session routing statistics in Switchyard using the --routing-stats-json flag. Analyze model calls, errors, and overhead for each logical session with detailed JSON output.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Switchyard Collects and Exposes Prometheus Metrics: A Deep Dive into the Rust Implementation](/NVIDIA-NeMo/Switchyard/how-switchyard-collect-expose-prometheus-metrics)

Learn how Switchyard collects and exposes Prometheus metrics using OpenTelemetry and an Axum endpoint. Discover its Rust implementation for histograms, counters, and gauges.

- Tags: deep-dive
- Published: 2026-08-21

### [What Is the Purpose of the switchyard-py Crate in NVIDIA Switchyard?](/NVIDIA-NeMo/Switchyard/what-is-purpose-switchyard-py-crate)

Discover the purpose of the switchyard-py crate. Access Switchyard's high-performance Rust server and routing algorithms directly from Python for programmatic control.

- Tags: internals
- Published: 2026-08-21

### [When to Use the Passthrough Algorithm in Switchyard: Complete Implementation Guide](/NVIDIA-NeMo/Switchyard/when-to-use-passthrough-algorithm-switchyard)

Learn when to use the Passthrough algorithm in Switchyard for static deployments, debugging, and performance baselines. Implement zero-overhead routing efficiently with this NVIDIA-NeMo guide.

- Tags: how-to-guide
- Published: 2026-08-21

### [How the Escalation Router Optimizes Costs in Switchyard](/NVIDIA-NeMo/Switchyard/how-does-escalation-router-switchyard-optimize-costs)

Discover how the Escalation router in NVIDIA Switchyard cuts LLM serving costs by intelligently routing conversations to affordable models, saving money while maintaining quality.

- Tags: performance
- Published: 2026-08-21

### [What Is Stage Routing in Switchyard? A Signal-Driven LLM Routing Algorithm](/NVIDIA-NeMo/Switchyard/what-is-stage-routing-switchyard-how-used)

Discover Stage routing in NVIDIA Switchyard, a signal-driven LLM algorithm that optimizes model selection for quality and cost based on real-time context and history. Learn how it works.

- Tags: deep-dive
- Published: 2026-08-21

### [How to Use the LLM Classifier Routing Algorithm in Switchyard](/NVIDIA-NeMo/Switchyard/how-to-use-llm-classifier-routing-algorithm-switchyard)

Master Switchyard's LLM Classifier routing algorithm for dynamic request routing. Optimize performance by intelligently directing tasks to weak or strong language models. Learn capability-based, escalation, and custom policies.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Does the Random Routing Algorithm in Switchyard Work?](/NVIDIA-NeMo/Switchyard/how-does-random-routing-algorithm-switchyard-work)

Discover how Switchyard's Random routing algorithm enables stateless traffic splitting. Learn about configurable weights and deterministic seeding without payload inspection.

- Tags: internals
- Published: 2026-08-21

### [How to Configure Switchyard's Routing Algorithms: TOML and Rust Integration Guide](/NVIDIA-NeMo/Switchyard/how-to-configure-switchyard-routing-algorithms)

Learn to configure Switchyard routing algorithms using TOML definitions and Rust integration. Explore stage_router and llm_classifier strategies with the switchyard.libsy Python package.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Switchyard Translates Streaming Responses Between Different LLM APIs](/NVIDIA-NeMo/Switchyard/how-switchyard-handles-streaming-responses-llm-apis)

Switchyard seamlessly translates streaming responses between LLM APIs by normalizing token streams into an intermediate representation and re-encoding them on-the-fly. No buffering required.

- Tags: how-to-guide
- Published: 2026-08-21

### [Provider-Neutral Request and Response Types in Switchyard: The Core Abstraction Layer](/NVIDIA-NeMo/Switchyard/what-are-provider-neutral-request-response-types-switchyard)

Explore provider-neutral request and response types in Switchyard. Learn how Request and Response structs abstract provider-specific API differences for seamless LLM integration.

- Tags: internals
- Published: 2026-08-21

### [How to Embed Switchyard's Routing Algorithms in a Custom Application](/NVIDIA-NeMo/Switchyard/how-to-embed-switchyard-routing-algorithms-custom-application)

Learn to embed Switchyard routing algorithms in your custom app. Integrate Switchyard's Rust engine via Python async iterators and handle CallModel events efficiently for LLM integration.

- Tags: how-to-guide
- Published: 2026-08-21

### [How the Switchyard Server Resolves LLM Routes: A Deep Dive into the Routing Pipeline](/NVIDIA-NeMo/Switchyard/how-does-switchyard-server-resolve-llm-routes)

Discover how the Switchyard server resolves LLM routes via a three-stage pipeline. Learn about identifier extraction, route lookup, and upstream target resolution from NVIDIA-NeMo/Switchyard.

- Tags: deep-dive
- Published: 2026-08-21

### [Switchyard Server HTTP Endpoints: Complete API Reference](/NVIDIA-NeMo/Switchyard/what-http-endpoints-switchyard-server-expose)

Explore Switchyard server HTTP endpoints including OpenAI chat, Anthropic messages, and Prometheus metrics. Discover the complete API reference for efficient integration.

- Tags: api-reference
- Published: 2026-08-21

### [How to Build and Run the Standalone Switchyard HTTP Proxy Server](/NVIDIA-NeMo/Switchyard/how-to-build-run-standalone-switchyard-http-proxy-server)

Learn to build and run the standalone Switchyard HTTP proxy server. Install via cargo or build from source, configure with TOML, and route OpenAI and Anthropic API requests using this native Rust binary.

- Tags: how-to-guide
- Published: 2026-08-21

### [Project Structure of the NVIDIA Switchyard Repository: Rust Core, Python Bindings, and Proxy Architecture](/NVIDIA-NeMo/Switchyard/what-is-project-structure-switchyard-repository)

Explore the NVIDIA Switchyard repository structure. Discover its Rust core for routing, Python bindings via PyO3, and proxy architecture. Learn how this mixed codebase powers efficient network solutions.

- Tags: architecture
- Published: 2026-08-21

### [How to Implement Content-Based Routing for LLMs with Switchyard](/NVIDIA-NeMo/Switchyard/how-to-implement-content-based-routing-llms-switchyard)

Implement content-based routing for LLMs with Switchyard. Inspect message payloads with the LLM Classifier and select backend models using textual patterns. Configure via Python or TOML.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Switchyard Optimizes LLM Costs Using Tiered Models](/NVIDIA-NeMo/Switchyard/can-switchyard-optimize-llm-costs-tiered-models)

Reduce LLM costs with Switchyard. Dynamically route requests between tiers using tiered models and the stage-router algorithm for efficient, cost-effective inference.

- Tags: how-to-guide
- Published: 2026-08-21

### [How to Use Switchyard for A/B Traffic Splitting with LLMs: A Complete Guide](/NVIDIA-NeMo/Switchyard/how-to-use-switchyard-for-llm-traffic-splitting)

Master A/B traffic splitting for LLMs with Switchyard. This guide shows you how to route requests between LLM backends effortlessly using zero-code configuration. Enhance your applications today.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Switchyard Translates Between OpenAI and Anthropic API Formats](/NVIDIA-NeMo/Switchyard/how-switchyard-translates-openai-anthropic-api-formats)

Discover how Switchyard enables seamless translation between OpenAI and Anthropic API formats. Leverage a unified interface for LLM development.

- Tags: how-to-guide
- Published: 2026-08-21

### [How to Set Up Switchyard as an LLM Traffic Orchestration Proxy](/NVIDIA-NeMo/Switchyard/how-to-set-up-switchyard-llm-traffic-orchestration-proxy)

Learn how to set up Switchyard, a Rust proxy, to orchestrate LLM traffic. Route requests to optimal backends with configurable algorithms like random split or LLM-as-classifier.

- Tags: how-to-guide
- Published: 2026-08-21

### [How Switchyard Exposes Codex-Compatible Model Metadata on the /v1/models Endpoint](/NVIDIA-NeMo/Switchyard/switchyard-codex-model-metadata-v1-models)

Learn how Switchyard exposes Codex-compatible model metadata via the /v1/models endpoint. Discover dual-format JSON payloads for OpenAI-style and Codex-specific model lists.

- Tags: api-reference
- Published: 2026-08-17

### [How to Integrate Switchyard with NVIDIA NIM or Ollama Backends](/NVIDIA-NeMo/Switchyard/integrate-switchyard-nvidia-nim-ollama)

Integrate Switchyard with NVIDIA NIM or Ollama backends easily. Learn how to configure routes.toml and environment variables for seamless OpenAI-compatible API calls.

- Tags: how-to-guide
- Published: 2026-08-17

### [How to Configure the Server Shutdown Timeout for Graceful Draining in NVIDIA Switchyard](/NVIDIA-NeMo/Switchyard/configure-server-shutdown-timeout-graceful-draining)

Configure the NVIDIA Switchyard server shutdown timeout for graceful draining. Override the default 30-second wait for in-flight requests using CLI flags, Rust structs, or Python parameters.

- Tags: how-to-guide
- Published: 2026-08-17

### [How Switchyard Handles Streaming Responses with Protocol Translation](/NVIDIA-NeMo/Switchyard/switchyard-streaming-responses-protocol-translation)

Learn how Switchyard handles streaming LLM responses with protocol translation. It decodes, translates, and re-encodes events for real-time format conversion.

- Tags: how-to-guide
- Published: 2026-08-17

### [How to Implement Custom Routing Algorithms in Switchyard's libsy Library](/NVIDIA-NeMo/Switchyard/implement-custom-routing-algorithms-libsy)

Implement custom routing algorithms in Switchyard's libsy library by creating Rust modules. Learn to compose algorithms and classifiers for model selection and telemetry.

- Tags: how-to-guide
- Published: 2026-08-17

### [How to Configure Multiple LLM Clients with Different Providers in One Switchyard Deployment](/NVIDIA-NeMo/Switchyard/configure-multiple-llm-clients-different-providers)

Learn to configure multiple LLM clients from providers like OpenAI Anthropic and NVIDIA in one Switchyard deployment Route requests dynamically using targets and stage routers

- Tags: how-to-guide
- Published: 2026-08-17

### [How Switchyard Handles Context Window Size Differences Between Models](/NVIDIA-NeMo/Switchyard/switchyard-handle-context-window-size-differences)

Switchyard prevents context window errors by validating requests against model limits, raising alerts before oversized prompts reach LLMs. Learn how Switchyard handles context window differences.

- Tags: internals
- Published: 2026-08-17

### [How to Validate a TOML Deployment Configuration Before Starting the Switchyard Server](/NVIDIA-NeMo/Switchyard/validate-toml-deployment-configuration)

Validate your TOML deployment configuration before starting the Switchyard server with the --dry-run flag. Detect and fix errors for clients targets and routes before deployment.

- Tags: how-to-guide
- Published: 2026-08-17

### [What Is the Difference Between switchyard-server and switchyard launch Command?](/NVIDIA-NeMo/Switchyard/difference-switchyard-server-launch-command)

Understand the difference between switchyard-server and switchyard launch. Learn how the Rust binary hosts the LLM-proxy API and the Python CLI orchestrates coding agents for tools like Claude Code.

- Tags: how-to-guide
- Published: 2026-08-17

### [How to Embed Switchyard Routing Algorithms in a Custom Rust Application](/NVIDIA-NeMo/Switchyard/embed-switchyard-routing-rust-application)

Embed Switchyard routing algorithms in your Rust app using switchyard-libsy. Integrate routing decisions directly into your service and manage HTTP traffic with ease.

- Tags: how-to-guide
- Published: 2026-08-17

### [How Switchyard Handles Upstream API Failures and Implements Fallback Logic](/NVIDIA-NeMo/Switchyard/switchyard-handle-upstream-api-failures-fallback)

Discover how Switchyard's two-layer resilience system manages upstream API failures. Learn about its client-side retry loop and server-side fallback for robust request routing.

- Tags: how-to-guide
- Published: 2026-08-17

### [How to Set Up TLS Encryption for the Switchyard Server](/NVIDIA-NeMo/Switchyard/setup-tls-encryption-switchyard-server)

Secure your Switchyard server with TLS encryption. Learn how to set up TLS using the cert and key flags for robust server security. Easy to follow guide.

- Tags: how-to-guide
- Published: 2026-08-17

### [Switchyard Protocol Translation Latency: How Much Overhead to Expect](/NVIDIA-NeMo/Switchyard/expected-latency-added-by-switchyard-protocol-translation-layer)

Discover the Switchyard protocol translation latency. Expect 1-3ms overhead per request by analyzing key Switchyard metrics. Optimize your network performance today.

- Tags: performance
- Published: 2026-08-15

### [How to Configure Authentication Keys for Upstream Providers in Switchyard](/NVIDIA-NeMo/Switchyard/configure-authentication-keys-for-upstream-providers-in-switchyard)

Learn how to configure upstream authentication keys in Switchyard. Reference environment variables in TOML for secure secret management and avoid embedding credentials.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Use NVIDIA Switchyard with Self-Hosted Models Like vLLM and Ollama](/NVIDIA-NeMo/Switchyard/can-switchyard-be-used-with-self-hosted-models-vllm-ollama)

Discover how to use NVIDIA Switchyard with self-hosted models like vLLM and Ollama. Route requests to OpenAI-compatible endpoints effortlessly with simple TOML configuration.

- Tags: how-to-guide
- Published: 2026-08-15

### [How the LLM Classifier in NVIDIA Switchyard Differentiates Between Strong and Weak Tiers](/NVIDIA-NeMo/Switchyard/how-llm-classifier-determines-strong-vs-weak-tiers)

Learn how the NVIDIA Switchyard LLM classifier routes requests between strong and weak tiers using a configurable solve probability threshold. Optimize your model's efficiency today.

- Tags: deep-dive
- Published: 2026-08-15

### [How to Debug Routing Decisions in Switchyard Using Logging and Tracing Features](/NVIDIA-NeMo/Switchyard/debug-routing-decisions-in-switchyard-with-logging-tracing)

Debug Switchyard routing decisions with Python logs and Prometheus counters. Enable SWITCHYARD_DEBUG=1 and query /metrics for detailed insights.

- Tags: how-to-guide
- Published: 2026-08-15

### [Passthrough vs Random Routing in Switchyard: A Complete Guide to LLM Request Routing Strategies](/NVIDIA-NeMo/Switchyard/difference-between-passthrough-and-random-routing-strategies)

Understand Switchyard's passthrough vs random routing strategies for LLM requests. Learn how to direct traffic efficiently to single or multiple models.

- Tags: deep-dive
- Published: 2026-08-15

### [How Switchyard Manages Concurrent Requests Across Different Routing Algorithms](/NVIDIA-NeMo/Switchyard/how-switchyard-handles-concurrent-requests-across-routing-algorithms)

Discover how Switchyard expertly handles concurrent requests with its Tokio async runtime and driver/channel architecture, enabling parallel routing algorithm execution for efficient request management.

- Tags: internals
- Published: 2026-08-15

### [How to Run the Native Rust Server Standalone Without the Python Launcher in Switchyard](/NVIDIA-NeMo/Switchyard/run-native-rust-server-standalone-without-python-launcher)

Learn to run Switchyard's native Rust server standalone without Python launchers. Build the switchyard server binary and use a TOML config for direct invocation.

- Tags: how-to-guide
- Published: 2026-08-15

### [What Is the Performance Overhead of Switchyard's Routing Decisions Compared to Direct API Calls?](/NVIDIA-NeMo/Switchyard/performance-overhead-of-switchyard-routing-decisions)

Discover Switchyard's routing overhead compared to direct API calls. Learn how its minimal latency impacts performance, typically under 5%.

- Tags: performance
- Published: 2026-08-15

### [How Switchyard Determines Which Upstream Format (OpenAI vs. Anthropic) to Use for a Target](/NVIDIA-NeMo/Switchyard/how-switchyard-selects-upstream-format-openai-vs-anthropic)

Learn how Switchyard selects upstream formats like OpenAI or Anthropic based on the wire api field in your provider configuration. Optimize your API integrations today.

- Tags: how-to-guide
- Published: 2026-08-15

### [Switchyard Server Health Check Endpoints: Complete 2024 Guide](/NVIDIA-NeMo/Switchyard/switchyard-server-health-check-endpoints)

Discover Switchyard server health check endpoints in our 2024 guide. Learn how to monitor your Switchyard server's operational status with the /health endpoint.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Switchyard Handles Context Window Exceeded Errors from Upstream LLM Providers](/NVIDIA-NeMo/Switchyard/how-switchyard-handles-context-window-exceeded-errors)

Learn how Switchyard effectively handles context window exceeded errors from upstream LLM providers, classifying and propagating these errors for robust application development.

- Tags: how-to-guide
- Published: 2026-08-15

### [Essential Fields Required in a Switchyard TOML Route Configuration](/NVIDIA-NeMo/Switchyard/required-fields-in-switchyard-toml-route-configuration)

Learn the essential fields for Switchyard TOML route configuration including id and type. Discover route type specific fields for efficient network routing. Maximize your Switchyard setup.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Switchyard Translates OpenAI Responses to Anthropic Messages Format](/NVIDIA-NeMo/Switchyard/how-switchyard-translates-openai-responses-to-anthropic-messages)

Learn how Switchyard translates OpenAI responses to Anthropic Messages format using its codec-based engine. Discover seamless integration and data transformation.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Configure Multiple Backend Targets for A/B Testing with Random Routing in Switchyard](/NVIDIA-NeMo/Switchyard/configure-multiple-backends-for-a-b-testing-with-random-routing)

Learn to configure multiple backend targets for A/B testing with random routing in Switchyard. Define targets, set route type to random, and control traffic splitting with weights and seeds.

- Tags: how-to-guide
- Published: 2026-08-15

### [Launcher Path vs Server Path in Switchyard: Architecture and Responsibilities Explained](/NVIDIA-NeMo/Switchyard/difference-between-launcher-path-and-server-path-in-switchyard)

Understand the Switchyard launcher path vs server path distinction. Learn how Python orchestration prepares your environment and Rust services execute LLM routing.

- Tags: architecture
- Published: 2026-08-15

### [How the Escalation Router Mode Functions in Switchyard's LLM Classifier Strategy](/NVIDIA-NeMo/Switchyard/how-escalation-router-mode-works-in-llm-classifier)

Discover how Switchyard's LLM classifier uses escalation router mode to handle tool signal failures and judge verdicts. Learn about configurable thresholds and contextual hand-offs.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Embed Switchyard's Routing Algorithms in a Rust Application Using switchyard-libsy](/NVIDIA-NeMo/Switchyard/embed-switchyard-routing-algorithms-in-rust-application-with-libsy)

Embed Switchyard's routing algorithms in your Rust application using switchyard-libsy. Orchestrate LLM traffic efficiently without HTTP dependencies and integrate seamlessly with any transport layer.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Switchyard Manages Fallback When a Model Backend Becomes Unavailable](/NVIDIA-NeMo/Switchyard/how-switchyard-handles-fallback-for-unavailable-backends)

Learn how Switchyard's multi-layered fallback strategy ensures model backend availability. Discover automatic retries and the First-Available algorithm with configurable routing policies.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Configure Custom Routing Algorithms in a Switchyard TOML Deployment File](/NVIDIA-NeMo/Switchyard/how-to-configure-custom-routing-algorithms-in-toml-deployment-file)

Learn to configure custom routing algorithms in Switchyard using TOML deployment files. Implement your own algorithms in Rust and register them for flexible network routing.

- Tags: how-to-guide
- Published: 2026-08-15

### [Stage‑Router vs LLM‑Classifier Routing in Switchyard: Key Differences and When to Use Each](/NVIDIA-NeMo/Switchyard/difference-between-stage-router-and-llm_classifier-routing-strategies)

Understand Stage-Router vs LLM-Classifier routing in Switchyard. Learn how each strategy balances quality and cost for efficient request handling. Discover key differences and use cases.

- Tags: deep-dive
- Published: 2026-08-15

### [How Switchyard Handles Streaming Responses Between OpenAI and Anthropic Formats](/NVIDIA-NeMo/Switchyard/how-does-switchyard-handle-streaming-responses-between-openai-and-anthropic-formats)

Learn how Switchyard effortlessly bridges OpenAI and Anthropic streaming formats by converting API responses into a neutral format for seamless interoperability.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Integrate Switchyard with Existing OpenAI SDK Client Applications](/NVIDIA-NeMo/Switchyard/integrate-switchyard-openai-sdk-client-applications)

Effortlessly integrate Switchyard with your OpenAI SDK apps. Launch the Rust server and set OPENAI_BASE_URL to localhost:4000/v1. No code modifications needed for seamless integration.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Manage Multiple Concurrent Routes and Targets in a Single Switchyard Deployment](/NVIDIA-NeMo/Switchyard/manage-multiple-concurrent-routes-targets-switchyard)

Learn to manage multiple concurrent routes and targets in a single Switchyard deployment. Define unlimited LLM clients and routes in one TOML file for efficient HTTP server operations.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Validate TOML Deployment Configuration with `switchyard-server --dry-run`](/NVIDIA-NeMo/Switchyard/validate-toml-deployment-config-switchyard-server-dry-run)

Validate TOML deployment configuration using switchyard-server --dry-run. Check your config for errors before deployment without starting the server.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Implement Custom Routing Algorithms in the libsy Crate: A Complete Guide](/NVIDIA-NeMo/Switchyard/implement-custom-routing-algorithms-libsy-crate)

Implement custom routing algorithms in the libsy crate by defining the Algorithm trait and using the Driver to manage decisions and model calls. Unlock advanced routing capabilities today.

- Tags: how-to-guide
- Published: 2026-08-15

### [Performance Optimization Strategies for High-Throughput LLM Traffic with Switchyard](/NVIDIA-NeMo/Switchyard/performance-optimization-high-throughput-llm-traffic-switchyard)

Boost LLM traffic performance with Switchyard. Implement Rust workers, keep-alive, batching, and health-aware routing for maximum throughput. Discover key strategies now.

- Tags: performance
- Published: 2026-08-15

### [PyO3 Python Bindings Architecture for switchyard_rust: A Deep Dive into the Three-Layer Design](/NVIDIA-NeMo/Switchyard/architecture-pyo3-python-bindings-switchyard_rust)

Explore the PyO3 Python bindings architecture for switchyard_rust. Discover its three-layer design featuring a Python loader, maturin-built extension, and Rust implementations for algorithms, models, and async streams.

- Tags: deep-dive
- Published: 2026-08-15

### [How to Debug Routing Failures Using Switchyard-Server Routing Log Files](/NVIDIA-NeMo/Switchyard/debug-routing-failures-switchyard-server-logs)

Debug routing failures in Switchyard-Server by enabling the routing log flag. Inspect log entries to pinpoint why LLM requests fail to route, identifying empty model fields or fallback reasons.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Implement Conversation Affinity and Turn-Based Routing in Switchyard](/NVIDIA-NeMo/Switchyard/implement-conversation-affinity-turn-based-routing-switchyard)

Implement conversation affinity and turn-based routing in Switchyard using AffinityRouter for sticky sessions and StageRouter for per-turn routing decisions. Learn how to enhance your conversational AI.

- Tags: how-to-guide
- Published: 2026-08-15

### [What Is the Purpose of the `switchyard-translation` Crate for Wire Format Conversion?](/NVIDIA-NeMo/Switchyard/purpose-switchyard-translation-crate-wire-format-conversion)

Discover how switchyard-translation converts LLM API wire formats like OpenAI and Anthropic into a neutral representation. Enable seamless cross-provider request handling in Switchyard.

- Tags: internals
- Published: 2026-08-15

### [How to Configure Environment Variable-Based API Key Management for LLM Providers in Switchyard](/NVIDIA-NeMo/Switchyard/configure-env-variable-api-key-management-llm-providers)

Learn to configure environment variable API key management for LLM providers in Switchyard. Securely manage keys using env_key field in launchers. Get started now.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Use the Switchyard Launch Command with Claude Code, Codex, and OpenClaw](/NVIDIA-NeMo/Switchyard/use-switchyard-launch-command-claude-codex-openclaw)

Master the switchyard launch command to integrate Claude Code, Codex, and OpenClaw. This guide explains how to spin up a native Switchyard server and route LLM requests efficiently.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Switchyard Handles Context Window Exceeded Errors and Maps Them to SwitchyardError](/NVIDIA-NeMo/Switchyard/handle-context-window-exceeded-errors-switchyarderror)

Learn how Switchyard gracefully handles 'context window exceeded' LLM errors, transforming them into consistent SwitchyardError responses for seamless integration.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Connect Self-Hosted vLLM Model Servers to Switchyard](/NVIDIA-NeMo/Switchyard/connect-self-hosted-vllm-model-servers-switchyard)

Connect self-hosted vLLM model servers to Switchyard. Configure llm_client, target, and passthrough routes in a TOML file for seamless integration with OpenAI-compatible endpoints.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Implement Passthrough Routes for Direct Model Access in Switchyard](/NVIDIA-NeMo/Switchyard/implement-passthrough-routes-direct-model-access)

Learn to implement passthrough routes in Switchyard using type noop for direct LLM model access. Bypass routing logic and streamline your application for faster performance. Get direct access now!

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Set Up Prometheus Metrics Collection in Switchyard Server: Complete Configuration Guide](/NVIDIA-NeMo/Switchyard/setup-prometheus-metrics-collection-switchyard-server)

Learn how to set up Prometheus metrics collection in Switchyard server easily. Configure the metrics endpoint and start monitoring your server's performance with this quick guide.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Configure an Escalation Router with an LLM Judge for Automatic Model Tier Promotion in Switchyard](/NVIDIA-NeMo/Switchyard/configure-escalation-router-llm-judge-model-promotion)

Configure an escalation router with an LLM judge in Switchyard to automatically promote model tiers. Enable efficient models to hand off requests to capable models when needed.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Stage-Router Routing Uses Tool-Result Signals to Drive Agent-Based Routing in Switchyard](/NVIDIA-NeMo/Switchyard/stage-router-routing-agent-tool-result-signaling)

Learn how stage-router routing in Switchyard uses tool-result signals from agents to intelligently select between expensive and efficient models based on a confidence score derived from error severity, production intensity, and...

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Implement LLM-Based Classifier Routing for Intelligent Request Distribution in Switchyard](/NVIDIA-NeMo/Switchyard/implement-llm-classifier-routing-intelligent-distribution)

Implement LLM-based classifier routing in Switchyard. Use the llm_task_classifier algorithm for intelligent request distribution and content-aware traffic management across your model infrastructure.

- Tags: how-to-guide
- Published: 2026-08-15

### [Switchyard Standalone Server vs. Embedded Runtime: A Complete Deployment Guide](/NVIDIA-NeMo/Switchyard/switchyard-standalone-rust-proxy-vs-embedded-runtime)

Understand Switchyard standalone server vs embedded runtime. Learn how to deploy Switchyard as an independent HTTP proxy or integrate its routing core directly into your Rust applications.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Set Up TLS Encryption for Switchyard-Server Production Deployment](/NVIDIA-NeMo/Switchyard/setup-tls-encryption-switchyard-server-production)

Secure your Switchyard server production deployment with TLS encryption. Easily enable HTTPS by providing your TLS cert and key files to the binary for robust security.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Implement Health-Aware Fallback with Automatic Endpoint Failover in Switchyard](/NVIDIA-NeMo/Switchyard/implement-health-aware-fallback-endpoint-failover-switchyard)

Implement health-aware fallback and automatic endpoint failover in Switchyard. Learn how continuous health probes, scoring, and routing ensure traffic always goes to the best available endpoint.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Translate Between OpenAI Chat Completions and Anthropic Messages API Formats in Switchyard](/NVIDIA-NeMo/Switchyard/translate-openai-anthropic-api-formats-switchyard)

Effortlessly translate between OpenAI Chat Completions and Anthropic Messages API formats using Switchyard. Learn how its translation crate and codecs streamline your LLM integrations.

- Tags: how-to-guide
- Published: 2026-08-15

### [How to Configure Weighted Random Routing for A/B Testing in Switchyard TOML Deployments](/NVIDIA-NeMo/Switchyard/configure-weighted-random-routing-ab-testing-switchyard-toml)

Configure weighted random routing for A B testing in Switchyard TOML with weights array to split traffic proportionally across models for reproducible results.

- Tags: how-to-guide
- Published: 2026-08-15

### [How Switchyard Implements Multi-Backend Routing for LLM Traffic Orchestration](/NVIDIA-NeMo/Switchyard/how-switchyard-implements-multi-backend-routing-llm-traffic)

Discover how Switchyard implements multi-backend routing for LLM traffic orchestration using its Rust HTTP server for intelligent distribution across OpenAI, Anthropic, and NVIDIA NIM.

- Tags: how-to-guide
- Published: 2026-08-15

### [How TOML Configuration Is Loaded at Server Startup Using Environment Variables in Switchyard](/NVIDIA-NeMo/Switchyard/loading-toml-configuration-server-startup-environment-variables)

Learn how Switchyard loads TOML configuration at server startup. Discover how it reads paths, deserializes into ServerConfig, and injects environment variable values for runtime server state.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure Switchyard for OpenClaw Launcher Integration](/NVIDIA-NeMo/Switchyard/configuring-switchyard-openclaw-launcher-integration)

Easily integrate OpenClaw with Switchyard using the built-in launcher. Auto-generates workspaces and routes LLM traffic through Switchyard's proxy seamlessly. Get started now.

- Tags: how-to-guide
- Published: 2026-08-14

### [Switchyard Route Types Comparison: Passthrough, Random, LLM Classifier, and Stage Router Explained](/NVIDIA-NeMo/Switchyard/comparing-passthrough-random-llm_classifier-stage_router-route-types)

Explore Switchyard route types: passthrough, random, LLM classifier, and stage router. Understand how each strategy orchestrates traffic and classifies workloads for your applications.

- Tags: comparison
- Published: 2026-08-14

### [PyO3 Bindings in Switchyard-Py: How Python Calls Rust in NVIDIA's Network Routing Framework](/NVIDIA-NeMo/Switchyard/exploring-pyo3-bindings-switchyard-py)

Explore PyO3 bindings in switchyard-py. Discover how Python code seamlessly calls Switchyard's Rust core for high-performance networking via a Python-friendly API.

- Tags: deep-dive
- Published: 2026-08-14

### [How Model IDs and Route IDs Work in Switchyard Request Routing](/NVIDIA-NeMo/Switchyard/how-model-ids-route-ids-work-request-routing)

Understand how Model IDs and Route IDs in Switchyard enable efficient LLM request routing. Learn how Switchyard selects optimal backends for your requests.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Set Up Self-Hosted vLLM Models with Switchyard: A Complete Guide](/NVIDIA-NeMo/Switchyard/setting-up-self-hosted-vllm-models-switchyard)

Set up self-hosted vLLM models with Switchyard. Install the CLI, configure your TOML file, and run Switchyard as a proxy to route OpenAI requests to local vLLM endpoints.

- Tags: how-to-guide
- Published: 2026-08-14

### [How Switchyard Processes Streaming Responses and Events: A Deep Dive into NVIDIA's LLM Routing Architecture](/NVIDIA-NeMo/Switchyard/processing-streaming-responses-events-switchyard)

Learn how Switchyard processes streaming responses and events with its three-layer architecture for efficient LLM routing. Discover NVIDIA's innovative approach.

- Tags: deep-dive
- Published: 2026-08-14

### [How to Configure Random Routing Weights for A/B Testing in Switchyard](/NVIDIA-NeMo/Switchyard/understanding-random-routing-weights-ab-testing)

Configure random routing weights for A/B testing in Switchyard. Split traffic proportionally across model targets using relative weight values for simple A/B testing.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Integrate the Codex CLI with the Switchyard Launcher](/NVIDIA-NeMo/Switchyard/integrating-codex-cli-switchyard-launcher)

Learn to integrate the Codex CLI with the Switchyard launcher. Follow steps to launch a native server, generate a model catalog, and inject provider overrides for local proxy endpoint access.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Deploy Switchyard as a Standalone HTTP Proxy: Complete Setup Guide](/NVIDIA-NeMo/Switchyard/deploying-switchyard-standalone-http-proxy)

Deploy Switchyard as a standalone HTTP proxy with our complete guide. Learn to install the server, configure routes, and launch with simple CLI flags for efficient proxy management.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Implement an Escalation Router with LLM Judge Detection in Switchyard (3-Stage Pipeline)](/NVIDIA-NeMo/Switchyard/implementing-escalation-router-llm-judge-detection)

Implement an Escalation Router with LLM Judge detection in Switchyard. Chain LLM stages and use decision functions for model escalation. Learn more from NVIDIA-NeMo.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure Retry Policies for Upstream LLM Calls in Switchyard](/NVIDIA-NeMo/Switchyard/configuring-retry-policies-upstream-llm-calls)

Learn to configure retry policies for upstream LLM calls in Switchyard using the max_retries field. Benefit from automatic exponential back-off and Retry-After header support.

- Tags: how-to-guide
- Published: 2026-08-14

### [How Switchyard Handles Context Window Exceeded Errors and Eviction Policies in LLM Routing](/NVIDIA-NeMo/Switchyard/handling-context-window-exceeded-errors-eviction-policies)

Switchyard handles context window exceeded errors by triggering evictions and routing requests away from failing backends for robust LLM traffic orchestration.

- Tags: deep-dive
- Published: 2026-08-14

### [How to Monitor the Switchyard Server with Exposed Prometheus Metrics](/NVIDIA-NeMo/Switchyard/monitoring-switchyard-server-exposed-metrics)

Monitor your Switchyard server easily with Prometheus metrics. Access operational data via the unauthenticated /metrics HTTP endpoint. Learn how now.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Use Stage-Router Routing with Tool-Result and Progress Signals in Switchyard](/NVIDIA-NeMo/Switchyard/using-stage-router-routing-tool-result-progress-signals)

Learn how to use Stage-Router with tool-result and progress signals in Switchyard. Automatically route agent requests between model tiers based on tool output analysis.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Secure API Keys with Environment Variables in Switchyard](/NVIDIA-NeMo/Switchyard/securing-api-keys-environment-variables-switchyard)

Secure API keys in Switchyard by using environment variables. Learn how to keep secrets out of source control and configuration files for enhanced security.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Integrate Claude Code with the Switchyard Launcher: A Complete Guide](/NVIDIA-NeMo/Switchyard/integrating-claude-code-switchyard-launcher)

Integrate Claude code with the Switchyard launcher easily. This guide shows how to use the Rust server, configure environment variables, and serve a live statistics UI for your Claude binary.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Translate Between OpenAI Chat, OpenAI Responses, and Anthropic Messages Formats in Switchyard](/NVIDIA-NeMo/Switchyard/translating-between-openai-chat-responses-anthropic-messages-formats)

Learn to translate between OpenAI Chat, OpenAI Responses, and Anthropic Messages formats using Switchyard's TranslationEngine. Seamlessly convert request and response payloads.

- Tags: how-to-guide
- Published: 2026-08-14

### [How Provider-Neutral Protocol Types Work in Switchyard-Protocol: A Deep Dive into NVIDIA's LLM Abstraction Layer](/NVIDIA-NeMo/Switchyard/how-provider-neutral-protocol-types-work-switchyard-protocol)

Learn how provider-neutral protocol types in NVIDIA Switchyard abstract LLM conversations. Route and translate requests across providers like OpenAI and Anthropic without vendor lock-in.

- Tags: deep-dive
- Published: 2026-08-14

### [How to Set Up Passthrough Routes for Direct Model Access in Switchyard](/NVIDIA-NeMo/Switchyard/setting-up-passthrough-routes-direct-model-access)

Learn to set up passthrough routes in Switchyard for direct model access. Achieve minimal latency with this proxy to upstream model providers. Get started now!

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Implement LLM Classifier Routing for Intelligent Model Selection in Switchyard](/NVIDIA-NeMo/Switchyard/implementing-llm-classifier-routing-intelligent-model-selection)

Implement LLM classifier routing in NVIDIA Switchyard with libsy. Route requests to judge or target models based on confidence for efficient, intelligent model selection.

- Tags: how-to-guide
- Published: 2026-08-14

### [How to Configure TOML Deployment Files for a Switchyard Server](/NVIDIA-NeMo/Switchyard/configuring-toml-deployment-files-switchyard-server)

Learn to configure TOML deployment files for your Switchyard server. Define routes, LLM endpoints, and runtime behavior efficiently using structured tables in this quick guide.

- Tags: how-to-guide
- Published: 2026-08-14

### [Switchyard Three-Layer Architecture Explained: LLM Clients, Targets, and Routes](/NVIDIA-NeMo/Switchyard/understanding-switchyard-three-layer-architecture-clients-targets-routes)

Explore Switchyard's three-layer architecture: LLM clients, targets, and routes. Understand how this system efficiently manages LLM routing and model selection for your applications.

- Tags: architecture
- Published: 2026-08-14

### [How Switchyard Routes LLM Traffic Between Multiple Model Providers: A Technical Deep Dive](/NVIDIA-NeMo/Switchyard/how-switchyard-routes-llm-traffic-between-multiple-model-providers)

Discover how Switchyard routes LLM traffic between providers using Rust algorithms like Stage-Router and LLM-Task-Classifier for intelligent request distribution based on tool signals and confidence.

- Tags: deep-dive
- Published: 2026-08-14

