RuntimeModels and Scope Separation in NVIDIA NeMo Switchyard: Isolating Parent and Sub-Agent Traffic
RuntimeModels in Switchyard are per-request dictionaries that map logical model categories to concrete model IDs, while Scope separation ensures parent and sub-agent traffic remain isolated by instantiating fresh routing contexts with independent RuntimeModels for every sub-agent invocation.
Switchyard is a Rust-based routing engine for LLM requests developed by NVIDIA. Understanding how it handles RuntimeModels and maintains strict isolation between parent agents and sub-agents through Scope separation is essential for building secure, multi-agent systems that prevent routing conflicts and quota leakage.
What Are RuntimeModels in Switchyard?
RuntimeModels are per-request configuration dictionaries that tell the routing algorithm which concrete model IDs are available for each logical category. Rather than hardcoding model names in the routing logic, Switchyard uses abstract buckets such as efficient, capable, and any.
The following example from the repository’s README.md (lines 30‑34) demonstrates the mapping structure:
runtime_models = {
"efficient": ["fast"],
"capable": ["quality"],
"any": ["quality", "fast"],
}
This abstraction provides two critical capabilities:
- Deployment flexibility: Different deployments can expose different concrete model IDs while maintaining consistent logical names across the codebase.
- Dynamic selection: The algorithm can query the caller for the specific model list valid for the current request, accommodating scenarios where only a subset of models (e.g., only "efficient" models) is available for a particular traffic pattern.
When the routing algorithm executes, it receives the request together with this mapping and yields a stream of Step objects (CallModel and Done). The logic then selects a model from the appropriate category list, invokes it, and can fall back to alternative candidates within the same category without restarting the request context.
How Scope Separation Maintains Isolation Between Parent and Sub-Agent Traffic
Switchyard supports sub-agents—child agents spawned by a parent LLM to perform sub-tasks such as tool-use calls. To prevent the parent’s routing decisions from contaminating the sub-agent’s traffic, Switchyard introduces scope separation, which treats each sub-agent as an independent routing entity.
The Parent and Sub-Agent Scope Model
Each execution context operates within a distinct Scope that encapsulates its routing state:
- Parent Scope: The main agent handling the original user request. It maintains its own
runtime_modelsmap and routing state throughout the request lifecycle. - Sub-Agent Scope: Any child agent invoked by the parent. It receives a new, independent scope containing a fresh copy of the
runtime_modelsdictionary and a dedicated routing algorithm instance.
Because each scope owns its own runtime model dictionary and routing state, the parent’s choice of "efficient" versus "capable" models does not affect the sub-agent’s routing decisions. The sub-agent’s traffic is routed as if it were a completely separate request, delivering three critical guarantees:
- Isolation: Model IDs selected for the parent are invisible to the sub-agent.
- Safety: A mis-routed sub-agent cannot inadvertently consume the parent’s quota or trigger a different routing policy.
- Correctness: The sub-agent can be configured with its own set of targets (e.g., a specialized tool model) without interfering with the parent’s routing logic.
Implementation in the Sub-Agent Routing Algorithm
The isolation mechanism is implemented in crates/libsy/src/algorithms/subagent.rs. This file contains the sub-agent-aware routing algorithm that creates a fresh Scope object for every sub-agent request and passes it to the underlying Algorithm::run_stream call.
According to the source code, the algorithm ensures that routing decisions for a sub-agent are made independently of the parent’s decision by instantiating a new scope boundary. The runtime model dictionaries are passed at execution time via crates/switchyard-nemo-relay-plugin/src/runtime.rs, which handles the dictionary injection into the algorithm stream.
Practical Implementation Example
The following Python driver pattern demonstrates how runtime_models is passed to run_stream and how Switchyard automatically scopes sub-agent traffic:
from switchyard.libsy import LlmResponse, Step
from switchyard.libsy.algorithms import stage_router
# Build the algorithm (stage router is one of the supported algorithms)
algorithm = stage_router(picker="efficient_first", confidence_threshold=0.5)
# Define the runtime model groups for the current request
runtime_models = {
"efficient": ["fast"],
"capable": ["quality"],
"any": ["quality", "fast"],
}
async def route(request: dict, clients: dict) -> LlmResponse.Agg | LlmResponse.Stream:
"""
Run a Switchyard algorithm with isolated scope.
Each sub-agent that the algorithm spawns will get its own copy
of `runtime_models`, guaranteeing traffic isolation.
"""
async for step in algorithm.run_stream(request, runtime_models):
match step:
case Step.CallModel(call):
# Call the model indicated by `call.models` (already scoped)
resp = await clients[call.models[0]].call(call.request)
call.respond(LlmResponse.Agg(resp))
case Step.Done(outcome):
# Final answer – either already produced or we must call the model
return outcome.response or await clients[outcome.selected_model_ids[0]].call(outcome.request)
raise RuntimeError("Algorithm finished without a decision")
In this implementation, the runtime_models dictionary is passed to run_stream. If the algorithm spawns a sub-agent (e.g., via a tool call), Switchyard creates a new scope with a copy of this dictionary. As documented in docs/routing_algorithms/subagent_routing.md, this design ensures that the sub-agent’s routing decisions remain completely isolated from the parent’s context.
Summary
- RuntimeModels provide per-request logical-to-physical model mappings via dictionaries that abstract concrete model IDs into categories like "efficient" and "capable".
- Scope separation instantiates fresh routing contexts for sub-agents, preventing traffic contamination between parent and child agents.
- The sub-agent routing algorithm in
crates/libsy/src/algorithms/subagent.rsenforces isolation by creating independentScopeobjects with dedicatedAlgorithm::run_streaminstances. - Each scope owns its own
runtime_modelsdictionary and routing state, ensuring sub-agents cannot access parent quotas or routing policies. - This architecture enables safe multi-agent deployments where specialized sub-agents can operate with distinct model configurations without interfering with parent logic.
Frequently Asked Questions
What is the exact format of the RuntimeModels dictionary in Switchyard?
The RuntimeModels dictionary is a standard key-value mapping where string keys represent logical categories (e.g., "efficient", "capable", "any") and values are lists of concrete model ID strings. This format allows the routing algorithm in crates/libsy/src/algorithms/ to dynamically select from available deployments without hardcoding model names.
How does Switchyard prevent a sub-agent from affecting the parent's routing state?
Switchyard prevents contamination by creating a new Scope object with a fresh copy of the runtime_models dictionary for every sub-agent invocation. As implemented in crates/libsy/src/algorithms/subagent.rs, the sub-agent receives a completely independent routing context, ensuring the parent’s model selections, quota counters, and policy state remain invisible to the child agent.
Can sub-agents use different model configurations than their parent agent?
Yes. Because each sub-agent receives its own runtime_models mapping through scope separation, you can configure specialized tool models (e.g., "code-interpreter") for sub-agents while the parent uses general-purpose models. Since the dictionaries are isolated copies, the parent’s routing logic will never conflate its "efficient" or "capable" selections with the sub-agent’s distinct target list.
Where is the Scope separation logic located in the Switchyard codebase?
The core Scope separation logic resides in crates/libsy/src/algorithms/subagent.rs, which handles sub-agent-aware routing by spawning fresh Scope instances. The runtime model dictionary passing occurs in crates/switchyard-nemo-relay-plugin/src/runtime.rs, while the high-level design is documented in docs/routing_algorithms/subagent_routing.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →