# How to Monitor Per-Session Routing Statistics in Switchyard

> Monitor per-session routing statistics in Switchyard using the --routing-stats-json flag. Analyze model calls, errors, and overhead for each logical session with detailed JSON output.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-08-21

---

**Switchyard emits detailed per-session routing counters as JSON when you run benchmarks with the `--routing-stats-json` flag, enabling precise analysis of model calls, errors, and routing overhead for every logical session.**

Monitoring per-session routing statistics in Switchyard allows you to diagnose performance bottlenecks and track error rates across individual conversation sessions in the NVIDIA-NeMo/Switchyard routing engine. By capturing granular metrics such as model invocation counts and classifier decision paths, you gain visibility into how the Rust-based routing layer handles traffic at the session level. This guide explains how to enable, extract, and analyze these statistics using the built-in CLI flags and Python APIs.

## Enabling Per-Session Routing Statistics Output

To capture routing statistics, invoke the Switchyard soak test runner with the `--routing-stats-json` flag followed by a target file path. The optional `--routing-stats-status` flag indicates whether the statistics file was successfully produced during the run.

```bash
switchyard-soak \
    --config routes.toml \
    --concurrency 64 \
    --request-count 5000 \
    --routing-stats-json ./output/routing_stats.json \
    --routing-stats-status present

```

When the run completes, the specified JSON file contains a snapshot of all counters aggregated by session ID. The `switchyard-soak` CLI entry point in [`scripts/run_local_soak_test.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/scripts/run_local_soak_test.py) handles argument parsing and passes these paths to the underlying benchmark harness.

## Understanding the JSON Output Format

The generated JSON document uses a flat mapping keyed by session identifier. Each session entry records discrete counters maintained by the Rust routing engine throughout the request lifecycle.

```json
{
  "sessions": {
    "session-01": {
      "model_calls": 12,
      "model_errors": 0,
      "classifier_calls": 3,
      "classifier_errors": 0,
      "routing_overhead_ms": 5.4
    },
    "session-02": {
      "model_calls": 8,
      "model_errors": 1,
      "classifier_calls": 2,
      "classifier_errors": 0,
      "routing_overhead_ms": 3.2
    }
  }
}

```

Key metrics include **model_calls** (total LLM invocations), **model_errors** (failed model requests), **classifier_calls** (routing decision evaluations), and **routing_overhead_ms** (time spent in routing logic excluding model inference).

## How Routing Statistics Are Collected Internally

The per-session monitoring pipeline spans both the Rust core and the Python benchmark framework.

### Rust Engine Counter Collection

The Rust routing engine in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) maintains in-memory counters for every active session. As each request belonging to a session traverses the routing graph, the engine increments the appropriate counters for model invocations, classifier evaluations, and error conditions without blocking the hot path.

### Python Snapshot Persistence

When a benchmark run finishes, [`benchmark/run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run_manifest.py) invokes `RoutingSnapshot.capture()` via the Rust bindings to serialize the current counter state. This method extracts the in-memory statistics from the Rust heap and returns them to Python, where the helper writes the snapshot to the path supplied by `--routing-stats-json` and registers the file reference in the run manifest for downstream reporting.

## Parsing and Analyzing Routing Statistics

You can process the generated JSON directly or access it through the run manifest object.

### Direct File Parsing

Use standard Python libraries to load and iterate over session statistics:

```python
import json
from pathlib import Path

stats_path = Path("./output/routing_stats.json")
stats = json.loads(stats_path.read_text())

for session_id, data in stats["sessions"].items():
    print(f"Session {session_id}:")
    print(f"  Model calls       : {data['model_calls']}")
    print(f"  Model errors      : {data['model_errors']}")
    print(f"  Classifier calls  : {data['classifier_calls']}")
    print(f"  Routing overhead  : {data['routing_overhead_ms']} ms")

```

### Accessing via Run Manifest

For integrated benchmark reports, load the statistics from the manifest written by [`run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/run_manifest.py):

```python
from benchmark.run_manifest import load_manifest

manifest = load_manifest(Path("./output/manifest.json"))
routing_stats = manifest["outcomes"]["routing_stats_json"]

# routing_stats contains the same structure as the direct JSON file

total_model_calls = sum(
    s["model_calls"] for s in routing_stats["sessions"].values()
)

```

## Summary

- **Enable collection** by passing `--routing-stats-json <path>` when running `switchyard-soak` to generate a JSON snapshot of per-session counters.
- **Rust engine metrics** are tracked in memory by [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py) and include model calls, errors, classifier decisions, and routing overhead timing.
- **Snapshot serialization** occurs via `RoutingSnapshot.capture()` in [`benchmark/run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run_manifest.py), which persists data at the end of each benchmark run.
- **Consume the data** either by reading the JSON file directly or loading it from the run manifest using the `load_manifest()` helper.

## Frequently Asked Questions

### What flags are required to capture per-session routing statistics in Switchyard?

You must provide the `--routing-stats-json` flag followed by a file path where the JSON statistics should be written. Optionally, add `--routing-stats-status` to indicate whether the statistics file was produced, which is useful for automated benchmark orchestration that checks for output presence.

### Where does Switchyard store the routing counters while a session is active?

The Rust routing engine maintains counters in memory during request processing. According to the Switchyard source code in [`switchyard_rust/libsy.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/switchyard_rust/libsy.py), each session updates in-memory counters as requests flow through the routing graph, ensuring zero I/O overhead on the hot path until the final snapshot is captured.

### How can I validate that routing statistics are correctly integrated into my benchmark results?

The test suite in [`tests/test_run_manifest.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tests/test_run_manifest.py) validates JSON creation and manifest integration. You can verify your own runs by checking that the `routing_stats_json` key exists in the manifest outcomes dictionary after calling `load_manifest()`, or by asserting that the `--routing-stats-json` file contains the expected `sessions` top-level key with numeric counter values.