# RuntimeModels and Scope Separation in NVIDIA NeMo Switchyard: Isolating Parent and Sub-Agent Traffic

> Discover how Switchyard's RuntimeModels and Scope separation isolate parent and sub-agent traffic. Learn about per-request dictionaries and independent routing contexts for robust agent communication.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: deep-dive
- Published: 2026-09-13

---

**RuntimeModels in Switchyard are per-request dictionaries that map logical model categories to concrete model IDs, while Scope separation ensures parent and sub-agent traffic remain isolated by instantiating fresh routing contexts with independent RuntimeModels for every sub-agent invocation.**

Switchyard is a Rust-based routing engine for LLM requests developed by NVIDIA. Understanding how it handles **RuntimeModels** and maintains strict isolation between parent agents and sub-agents through **Scope separation** is essential for building secure, multi-agent systems that prevent routing conflicts and quota leakage.

## What Are RuntimeModels in Switchyard?

**RuntimeModels** are per-request configuration dictionaries that tell the routing algorithm which concrete model IDs are available for each logical category. Rather than hardcoding model names in the routing logic, Switchyard uses abstract buckets such as *efficient*, *capable*, and *any*.

The following example from the repository’s [`README.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/README.md) (lines 30‑34) demonstrates the mapping structure:

```python
runtime_models = {
    "efficient": ["fast"],
    "capable":  ["quality"],
    "any":       ["quality", "fast"],
}

```

This abstraction provides two critical capabilities:

- **Deployment flexibility:** Different deployments can expose different concrete model IDs while maintaining consistent logical names across the codebase.
- **Dynamic selection:** The algorithm can query the caller for the specific model list valid for the current request, accommodating scenarios where only a subset of models (e.g., only "efficient" models) is available for a particular traffic pattern.

When the routing algorithm executes, it receives the request together with this mapping and yields a stream of `Step` objects (`CallModel` and `Done`). The logic then selects a model from the appropriate category list, invokes it, and can fall back to alternative candidates within the same category without restarting the request context.

## How Scope Separation Maintains Isolation Between Parent and Sub-Agent Traffic

Switchyard supports **sub-agents**—child agents spawned by a parent LLM to perform sub-tasks such as tool-use calls. To prevent the parent’s routing decisions from contaminating the sub-agent’s traffic, Switchyard introduces **scope separation**, which treats each sub-agent as an independent routing entity.

### The Parent and Sub-Agent Scope Model

Each execution context operates within a distinct **Scope** that encapsulates its routing state:

- **Parent Scope:** The main agent handling the original user request. It maintains its own `runtime_models` map and routing state throughout the request lifecycle.
- **Sub-Agent Scope:** Any child agent invoked by the parent. It receives a **new, independent scope** containing a fresh copy of the `runtime_models` dictionary and a dedicated routing algorithm instance.

Because each scope owns its own runtime model dictionary and routing state, the parent’s choice of "efficient" versus "capable" models does **not** affect the sub-agent’s routing decisions. The sub-agent’s traffic is routed as if it were a completely separate request, delivering three critical guarantees:

1. **Isolation:** Model IDs selected for the parent are invisible to the sub-agent.
2. **Safety:** A mis-routed sub-agent cannot inadvertently consume the parent’s quota or trigger a different routing policy.
3. **Correctness:** The sub-agent can be configured with its own set of targets (e.g., a specialized tool model) without interfering with the parent’s routing logic.

### Implementation in the Sub-Agent Routing Algorithm

The isolation mechanism is implemented in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs). This file contains the sub-agent-aware routing algorithm that creates a fresh `Scope` object for every sub-agent request and passes it to the underlying `Algorithm::run_stream` call.

According to the source code, the algorithm ensures that routing decisions for a sub-agent are made **independently** of the parent’s decision by instantiating a new scope boundary. The runtime model dictionaries are passed at execution time via [`crates/switchyard-nemo-relay-plugin/src/runtime.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/runtime.rs), which handles the dictionary injection into the algorithm stream.

### Practical Implementation Example

The following Python driver pattern demonstrates how `runtime_models` is passed to `run_stream` and how Switchyard automatically scopes sub-agent traffic:

```python
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router

# Build the algorithm (stage router is one of the supported algorithms)

algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)

# Define the runtime model groups for the current request

runtime_models = {
    "efficient": ["fast"],
    "capable":  ["quality"],
    "any":       ["quality", "fast"],
}

async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
    """
    Run a Switchyard algorithm with isolated scope.
    Each sub-agent that the algorithm spawns will get its own copy
    of `runtime_models`, guaranteeing traffic isolation.
    """
    async for step in algorithm.run_stream(request, runtime_models):
        match step:
            case Step.CallModel(call):
                # Call the model indicated by `call.models` (already scoped)

                resp = await clients[call.models[0]].call(call.request)
                call.respond(LlmResponse.Agg(resp))
            case Step.Done(outcome):
                # Final answer – either already produced or we must call the model

                return outcome.response or await clients[outcome.selected_model_ids[0]].call(outcome.request)
    raise RuntimeError("Algorithm finished without a decision")

```

In this implementation, the `runtime_models` dictionary is passed to `run_stream`. If the algorithm spawns a sub-agent (e.g., via a tool call), Switchyard creates a *new scope* with a **copy** of this dictionary. As documented in [`docs/routing_algorithms/subagent_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/subagent_routing.md), this design ensures that the sub-agent’s routing decisions remain completely isolated from the parent’s context.

## Summary

- **RuntimeModels** provide per-request logical-to-physical model mappings via dictionaries that abstract concrete model IDs into categories like "efficient" and "capable".
- **Scope separation** instantiates fresh routing contexts for sub-agents, preventing traffic contamination between parent and child agents.
- The **sub-agent routing algorithm** in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs) enforces isolation by creating independent `Scope` objects with dedicated `Algorithm::run_stream` instances.
- Each scope owns its own `runtime_models` dictionary and routing state, ensuring sub-agents cannot access parent quotas or routing policies.
- This architecture enables safe multi-agent deployments where specialized sub-agents can operate with distinct model configurations without interfering with parent logic.

## Frequently Asked Questions

### What is the exact format of the RuntimeModels dictionary in Switchyard?

The **RuntimeModels** dictionary is a standard key-value mapping where string keys represent logical categories (e.g., `"efficient"`, `"capable"`, `"any"`) and values are lists of concrete model ID strings. This format allows the routing algorithm in `crates/libsy/src/algorithms/` to dynamically select from available deployments without hardcoding model names.

### How does Switchyard prevent a sub-agent from affecting the parent's routing state?

Switchyard prevents contamination by creating a **new Scope object** with a fresh copy of the `runtime_models` dictionary for every sub-agent invocation. As implemented in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs), the sub-agent receives a completely independent routing context, ensuring the parent’s model selections, quota counters, and policy state remain invisible to the child agent.

### Can sub-agents use different model configurations than their parent agent?

Yes. Because each sub-agent receives its own `runtime_models` mapping through scope separation, you can configure specialized tool models (e.g., `"code-interpreter"`) for sub-agents while the parent uses general-purpose models. Since the dictionaries are isolated copies, the parent’s routing logic will never conflate its "efficient" or "capable" selections with the sub-agent’s distinct target list.

### Where is the Scope separation logic located in the Switchyard codebase?

The core **Scope separation** logic resides in [`crates/libsy/src/algorithms/subagent.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/libsy/src/algorithms/subagent.rs), which handles sub-agent-aware routing by spawning fresh `Scope` instances. The runtime model dictionary passing occurs in [`crates/switchyard-nemo-relay-plugin/src/runtime.rs`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/crates/switchyard-nemo-relay-plugin/src/runtime.rs), while the high-level design is documented in [`docs/routing_algorithms/subagent_routing.md`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/docs/routing_algorithms/subagent_routing.md).