How to Monitor Per-Session Routing Statistics in Switchyard
Switchyard emits detailed per-session routing counters as JSON when you run benchmarks with the --routing-stats-json flag, enabling precise analysis of model calls, errors, and routing overhead for every logical session.
Monitoring per-session routing statistics in Switchyard allows you to diagnose performance bottlenecks and track error rates across individual conversation sessions in the NVIDIA-NeMo/Switchyard routing engine. By capturing granular metrics such as model invocation counts and classifier decision paths, you gain visibility into how the Rust-based routing layer handles traffic at the session level. This guide explains how to enable, extract, and analyze these statistics using the built-in CLI flags and Python APIs.
Enabling Per-Session Routing Statistics Output
To capture routing statistics, invoke the Switchyard soak test runner with the --routing-stats-json flag followed by a target file path. The optional --routing-stats-status flag indicates whether the statistics file was successfully produced during the run.
switchyard-soak \
--config routes.toml \
--concurrency 64 \
--request-count 5000 \
--routing-stats-json ./output/routing_stats.json \
--routing-stats-status present
When the run completes, the specified JSON file contains a snapshot of all counters aggregated by session ID. The switchyard-soak CLI entry point in scripts/run_local_soak_test.py handles argument parsing and passes these paths to the underlying benchmark harness.
Understanding the JSON Output Format
The generated JSON document uses a flat mapping keyed by session identifier. Each session entry records discrete counters maintained by the Rust routing engine throughout the request lifecycle.
{
"sessions": {
"session-01": {
"model_calls": 12,
"model_errors": 0,
"classifier_calls": 3,
"classifier_errors": 0,
"routing_overhead_ms": 5.4
},
"session-02": {
"model_calls": 8,
"model_errors": 1,
"classifier_calls": 2,
"classifier_errors": 0,
"routing_overhead_ms": 3.2
}
}
}
Key metrics include model_calls (total LLM invocations), model_errors (failed model requests), classifier_calls (routing decision evaluations), and routing_overhead_ms (time spent in routing logic excluding model inference).
How Routing Statistics Are Collected Internally
The per-session monitoring pipeline spans both the Rust core and the Python benchmark framework.
Rust Engine Counter Collection
The Rust routing engine in switchyard_rust/libsy.py maintains in-memory counters for every active session. As each request belonging to a session traverses the routing graph, the engine increments the appropriate counters for model invocations, classifier evaluations, and error conditions without blocking the hot path.
Python Snapshot Persistence
When a benchmark run finishes, benchmark/run_manifest.py invokes RoutingSnapshot.capture() via the Rust bindings to serialize the current counter state. This method extracts the in-memory statistics from the Rust heap and returns them to Python, where the helper writes the snapshot to the path supplied by --routing-stats-json and registers the file reference in the run manifest for downstream reporting.
Parsing and Analyzing Routing Statistics
You can process the generated JSON directly or access it through the run manifest object.
Direct File Parsing
Use standard Python libraries to load and iterate over session statistics:
import json
from pathlib import Path
stats_path = Path("./output/routing_stats.json")
stats = json.loads(stats_path.read_text())
for session_id, data in stats["sessions"].items():
print(f"Session {session_id}:")
print(f" Model calls : {data['model_calls']}")
print(f" Model errors : {data['model_errors']}")
print(f" Classifier calls : {data['classifier_calls']}")
print(f" Routing overhead : {data['routing_overhead_ms']} ms")
Accessing via Run Manifest
For integrated benchmark reports, load the statistics from the manifest written by run_manifest.py:
from benchmark.run_manifest import load_manifest
manifest = load_manifest(Path("./output/manifest.json"))
routing_stats = manifest["outcomes"]["routing_stats_json"]
# routing_stats contains the same structure as the direct JSON file
total_model_calls = sum(
s["model_calls"] for s in routing_stats["sessions"].values()
)
Summary
- Enable collection by passing
--routing-stats-json <path>when runningswitchyard-soakto generate a JSON snapshot of per-session counters. - Rust engine metrics are tracked in memory by
switchyard_rust/libsy.pyand include model calls, errors, classifier decisions, and routing overhead timing. - Snapshot serialization occurs via
RoutingSnapshot.capture()inbenchmark/run_manifest.py, which persists data at the end of each benchmark run. - Consume the data either by reading the JSON file directly or loading it from the run manifest using the
load_manifest()helper.
Frequently Asked Questions
What flags are required to capture per-session routing statistics in Switchyard?
You must provide the --routing-stats-json flag followed by a file path where the JSON statistics should be written. Optionally, add --routing-stats-status to indicate whether the statistics file was produced, which is useful for automated benchmark orchestration that checks for output presence.
Where does Switchyard store the routing counters while a session is active?
The Rust routing engine maintains counters in memory during request processing. According to the Switchyard source code in switchyard_rust/libsy.py, each session updates in-memory counters as requests flow through the routing graph, ensuring zero I/O overhead on the hot path until the final snapshot is captured.
How can I validate that routing statistics are correctly integrated into my benchmark results?
The test suite in tests/test_run_manifest.py validates JSON creation and manifest integration. You can verify your own runs by checking that the routing_stats_json key exists in the manifest outcomes dictionary after calling load_manifest(), or by asserting that the --routing-stats-json file contains the expected sessions top-level key with numeric counter values.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →