# How Switchyard Implements Sub-Agent-Aware Routing with Passthrough and StageRouter Components

> Discover how Switchyard enables sub-agent-aware routing with passthrough and StageRouter. Learn how to delegate child work to isolated model groups for efficient processing.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: internals
- Published: 2026-09-13

---

**Switchyard implements sub-agent-aware routing by embedding a `subagents` configuration table inside primary route definitions, which wraps the parent algorithm with a `SubagentRouter` that inspects request metadata to delegate child work to isolated model groups.**

Switchyard is an open-source routing layer developed by NVIDIA that orchestrates LLM traffic across heterogeneous model backends. Sub-agent-aware routing allows the framework to delegate specific workloads—such as tool-driven workers or child agents—to dedicated models while preserving the primary routing logic for standard traffic. This architecture enables complex multi-agent workflows without polluting the main request pipeline.

## Understanding Sub-Agent Routing in Switchyard

In Switchyard, a **sub-agent** represents a delegated unit of work, typically triggered by tool calls or agentic workflows that spawn child processes. The framework distinguishes these requests using metadata flags, allowing the router to apply specialized policies independent of the parent route.

The core mechanism relies on **composable algorithms**. A primary algorithm—either `Passthrough` or `StageRouter`—handles regular traffic, while a `SubagentRouter` wraps it to intercept sub-agent requests. This design keeps concerns separated: the parent algorithm optimizes for general query characteristics, while the sub-agent router manages delegated workloads according to distinct performance or capability requirements.

## Configuration: Defining Sub-Agent Routes in TOML

Route definitions in Switchyard use TOML configuration files. To enable sub-agent-aware routing, you add a `[routes.<name>.subagents]` table adjacent to the primary route configuration.

### Passthrough with Fixed Sub-Agent Target

The simplest configuration uses `passthrough` for both parent and child routing, directing sub-agent work to a single dedicated model:

```toml
[routes.chat]
type = "passthrough"
target = "model/strong"

[routes.chat.subagents]
type = "passthrough"
target = "model/weak"

```

This configuration ensures that standard requests hit the strong model, while any work flagged as sub-agent traffic routes to the weak model.

### StageRouter with Classifier-Based Sub-Agent Routing

For more complex scenarios, you can pair a `StageRouter` parent with an `LlmClassifier` for dynamic sub-agent routing:

```toml
[routes.task]
type = "stage_router"
capable_target = "model/capable"
efficient_target = "model/efficient"
picker = "efficient_first"

[routes.task.subagents]
type = "llm_classifier"
mode = "custom"
default_target = "capable"
models = { any = ["worker", "reviewer"], judge = ["judge"] }
response_schema = """
{
  "type":"object",
  "properties":{"target":{"type":"string","enum":["capable","efficient"]}},
  "required":["target"]
}
"""
policy = { type = "target_selector", selector = "/target" }

```

The sub-agent classifier applies only to delegated work, allowing fine-grained control over which models handle specific tool-driven tasks.

## Algorithm Construction and the SubagentRouter

The construction process begins in [`crates/switchyard-runner/src/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/algorithm.rs). Both `AlgorithmSpec::Passthrough` and `AlgorithmSpec::StageRouter` accept an optional `subagents: Option<SubagentRouteConfig>` field that captures the secondary configuration.

When building the route, the system first constructs the primary algorithm, then calls the helper function `attach_subagent_router`. This function checks if a sub-agent configuration exists and, if present, wraps the primary algorithm with a `SubagentRouter`:

```rust
// Simplified construction flow from algorithm.rs
let primary_algorithm = match spec {
    AlgorithmSpec::Passthrough { target, subagents } => {
        // Build passthrough logic
    },
    AlgorithmSpec::StageRouter { capable_target, efficient_target, subagents } => {
        // Build stage router logic
    }
};

// Wrap with sub-agent router if configuration exists
if let Some(subagent_config) = subagents {
    attach_subagent_router(primary_algorithm, subagent_config)
}

```

This wrapping pattern ensures that sub-agent logic is injected transparently without modifying the parent algorithm's internal state.

## Core Implementation: The SubagentRouter Algorithm

The heart of sub-agent-aware routing resides in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs). The `SubagentRouter` struct implements the `Algorithm` trait and provides the delegation logic.

### Configuration and Pipeline Construction

The `SubagentRouterConfig` structure defines three critical parameters:
- **classifier**: Determines how sub-agent requests are categorized
- **default_target**: The fallback category when classification fails or is skipped
- **classify_trigger**: Determines when classification occurs (every request or new session only)

During initialization, `SubagentRouter::new` constructs a `FallThrough` pipeline that applies the classifier, falls back to the default target, and explicitly rejects the `message_hash_fallback` flag. This rejection is mandatory because sub-agents must be identified by harness metadata rather than message content hashing.

### Routing Logic and Metadata Inspection

The `route` method implementation checks the `is_subagent_work` flag in the request metadata:

```rust
// From crates/libsy/src/algorithms/subagent.rs
impl Algorithm for SubagentRouter {
    async fn route(&self, request: Request, driver: Driver) -> Result<RoutingDecision> {
        if request.metadata.is_subagent_work {
            // Delegate to sub-agent pipeline with isolated driver context
            self.subagent.execute(driver.for_subagent()?, request).await
        } else {
            // Pass to parent algorithm
            self.parent.route(request, driver).await
        }
    }
}

```

This metadata check occurs early in the routing cycle, ensuring minimal latency overhead for standard requests while providing a dedicated path for sub-agent workloads.

## Model Grouping and Isolation

Switchyard maintains strict isolation between parent and sub-agent model groups to prevent category collisions. When `AlgorithmSpec::runtime_model_names` executes, it generates two distinct maps:

- **parent**: Contains model groups for the primary algorithm (e.g., `capable`, `efficient`, `any`)
- **subagent**: Created by `subagent_runtime_model_names`, this distinct map isolates sub-agent categories

This separation prevents ambiguous routing decisions where a sub-agent request might accidentally resolve to the parent's `any` category. The isolation ensures that sub-agent policies reference only explicitly defined sub-agent models, maintaining predictable behavior in multi-tenant deployments.

## Affinity Routing for Sub-Agent Sessions

When the `classify_trigger` is set to `"new_session"`, the `SubagentRouter` attaches an `AffinityRouter::for_subagents()` processor. This router maintains session state, remembering the child model selected for the first request in a session and forcing subsequent requests to the same model.

Session affinity is crucial for sub-agent workflows that require stateful consistency, such as tool-use chains where the same worker model must handle all steps of a specific task. The affinity router persists this selection independently of the parent's affinity logic, allowing parent and child sessions to evolve separately.

## Supported Sub-Agent Policies

Switchyard supports two primary sub-agent routing policies, configured via the `type` field in the subagents table.

### Passthrough Sub-Agent Routing

The **Passthrough** policy (`SubagentRouterConfig::fixed_target`) routes every sub-agent request to a single, predetermined model. This approach minimizes latency and configuration complexity when the sub-agent workload is homogeneous.

### LlmClassifier Sub-Agent Routing

The **LlmClassifier** policy enables dynamic routing based on request content. However, the implementation restricts this mode to `custom` classification only—the code explicitly rejects other classification modes to prevent unintended routing loops or model selection conflicts.

When using `llm_classifier`, the configuration accepts:
- **models**: A mapping of categories to model lists (isolated from parent categories)
- **response_schema**: JSON schema constraining the classifier's output
- **policy**: A target selector extracting the routing decision from the classifier response

This policy allows sophisticated routing where different sub-agent tasks (e.g., code generation vs. validation) target different model capabilities.

## Summary

- **Configuration-driven**: Sub-agent-aware routing uses a `subagents` table in TOML route definitions to specify delegated routing logic alongside primary `passthrough` or `stage_router` algorithms.
- **Metadata-based delegation**: The `SubagentRouter` in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs) inspects `request.metadata.is_subagent_work` to bifurcate traffic between parent and child pipelines.
- **Algorithm wrapping**: The `attach_subagent_router` helper in [`crates/switchyard-runner/src/algorithm.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-runner/src/algorithm.rs) composes the primary algorithm with a `SubagentRouter` when sub-agent configuration is present.
- **Model isolation**: Separate `parent` and `subagent` model maps prevent category collisions and ensure sub-agent requests resolve only to explicitly configured child models.
- **Session affinity**: The `new_session` classify trigger enables `AffinityRouter` binding, maintaining model consistency across multi-step sub-agent workflows.
- **Policy flexibility**: Support for both fixed-target passthrough and custom `LlmClassifier` policies accommodates simple delegation and complex dynamic routing scenarios.

## Frequently Asked Questions

### How does Switchyard distinguish between regular requests and sub-agent requests?

Switchyard checks the `is_subagent_work` boolean flag in the request metadata structure defined in [`crates/protocol/src/request.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/protocol/src/request.rs). When this flag is true, the `SubagentRouter` intercepts the request and routes it through the configured sub-agent pipeline rather than the parent algorithm.

### Can I use different classification modes for the parent router and sub-agent router?

Yes. The parent router and sub-agent router maintain independent configurations. However, for `LlmClassifier` sub-agent routes, Switchyard explicitly restricts the mode to `custom` only, rejecting standard classification modes to prevent routing conflicts. The parent router can use any supported mode (efficient-first, capable-first, etc.) regardless of the sub-agent configuration.

### What happens if a sub-agent configuration references a model category not defined in the subagents table?

The routing will fail safely due to model isolation. Sub-agent routers use a separate model map created by `subagent_runtime_model_names`, distinct from the parent's model groups. If you reference a category like `any` in the sub-agent configuration without defining it in the subagents `models` table, the router cannot resolve the target and will return an error rather than falling back to the parent's `any` category.

### Is session affinity maintained separately for parent and sub-agent traffic?

Yes. When `classify_trigger = "new_session"` is configured, the `SubagentRouter` instantiates `AffinityRouter::for_subagents()`, which maintains its own session-to-model mapping. This operates independently of any affinity routing configured for the parent algorithm, allowing parent requests to round-robin across models while their associated sub-agent work remains pinned to a specific worker model for the session duration.