# How to Reproduce Terminal-Bench 2.1 Benchmark Results for Switchyard and Configure Escalation Deployment

> Reproduce Terminal-Bench 2.1 benchmark results for Switchyard. Learn how to prepare the dataset, establish a baseline, and configure escalation deployment using the provided TOML file.

- Repository: [NVIDIA-NeMo/Switchyard](https://github.com/NVIDIA-NeMo/Switchyard)
- Tags: how-to-guide
- Published: 2026-09-13

---

**To reproduce the Terminal-Bench 2.1 benchmark numbers for NVIDIA-NeMo/Switchyard, run a three-phase workflow using [`prepare_harbor_dataset.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/prepare_harbor_dataset.py) for dataset preparation, establish a baseline cost profile without routing, then execute against the [`tb21-escalation-opus-glm-deepseek.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tb21-escalation-opus-glm-deepseek.toml) configuration which controls model escalation through confirmation thresholds and windowed conversation history.**

The Switchyard repository implements an intelligent routing layer for LLM agents evaluated against the Terminal-Bench 2.1 dataset. Reproducing the published accuracy and cost metrics requires specific orchestration through Harbor and proper configuration of the escalation deployment TOML that defines weak-to-strong tier promotion logic.

## Reproducing the Terminal-Bench 2.1 Benchmark

The official benchmark workflow consists of three distinct phases executed via scripts in the `benchmark/` directory.

### Step 1 – Prepare the Closed-Book Dataset

First, generate the closed-book proxy of the TB 2.1 dataset used for evaluation. According to the source code in [`benchmark/prepare_harbor_dataset.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/prepare_harbor_dataset.py), invoke the preparation script with:

```bash
uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dataset terminal-bench/terminal-bench-2-1 \
  --output-dir benchmark/datasets/terminal-bench-2-1-closed-book \
  --overwrite

```

This pins agent images and creates the task manifest required by Harbor for isolated execution.

### Step 2 – Run the Baseline Without Switchyard

Establish reference metrics by sending requests directly to the upstream provider. The [`benchmark/run-baseline.sh`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run-baseline.sh) script handles Docker orchestration and environment setup. Execute:

```bash
bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/terminal-bench-2-1-closed-book \
  --model openai/gpt-5.5 \
  --agent codex \
  --reasoning-effort xhigh \
  --n-tasks 1 \
  --n-concurrent 1 \
  --max-retries 0

```

This generates cost and latency baselines against which Switchyard routing is compared.

### Step 3 – Execute with Switchyard Escalation

Run the benchmark through the Switchyard server using the escalation TOML. The configuration enables tier promotion from weak to strong models based on runtime signals:

```bash
bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/terminal-bench-2-1-closed-book \
  --server-config benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml \
  --model switchyard \
  --agent codex \
  --reasoning-effort xhigh \
  --n-tasks 1 \
  --n-concurrent 1 \
  --max-retries 0

```

Post-execution, locate the routing statistics in `benchmark/tb_runs/<run-id>/routing_stats_final.json`. The reported metrics include the 71-76% accuracy figures achieved at 13-30% lower cost than the Opus 4.8 baseline.

## Anatomy of the Escalation Deployment TOML

The escalation behavior is defined in [`benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml) following the v0.2.0 server-config schema. The file structures routing logic, model tiers, and escalation triggers.

### Router and Classifier Settings

The `[router]` section declares the route identifier supplied to the CLI via `--model`. The `[classifier]` block assigns a lightweight model to determine initial tier placement:

- `router.id = "tb21_escalation"`
- `classifier.model = "google/gemini-3.5-flash"`

### Tier Definitions

The `[tiers]` table delineates the weak and strong model pools:

- **Weak tier**: `[{ model = "moonshotai/kimi-k2.7-code" }]` handles initial requests.
- **Strong tier**: `[{ model = "anthropic/claude-opus-4.7" }]` receives escalated tasks.

### Escalation Logic and Parameters

The `[escalation]` block contains the core reproduction parameters for Terminal-Bench 2.1:

- `confirmations = 2`: Requires two validation signals before promoting to the strong tier.
- `recent_turn_window = 28`: Examines the last 28 conversation turns for escalation criteria.
- `window_message_chars = 500`: Limits the message history sent to the strong tier to 500 characters.

Additionally, setting `only_on_wrong_signal_escalation = true` restricts escalation to cases where the classifier explicitly signals an error or stall condition.

### Metadata and Notification Strings

The `[notes]` section provides guidance strings injected into prompts during tier transitions:

- `escalation_note`: Alerts the strong model that it is recovering a task from a weaker tier.
- `deescalation_note`: Indicates successful completion by the capable tier.

## Practical Commands to Run the Benchmark

Combine preparation and execution with increased concurrency for faster iteration:

```bash

# Prepare dataset

uv run --no-sync python benchmark/prepare_harbor_dataset.py \
  --source-dataset terminal-bench/terminal-bench-2-1 \
  --output-dir benchmark/datasets/terminal-bench-2-1-closed-book \
  --overwrite

# Run with escalation routing

bash benchmark/run-baseline.sh \
  --harbor-path benchmark/datasets/terminal-bench-2-1-closed-book \
  --server-config benchmark/routing-profiles/tb21-escalation-opus-glm-deepseek.toml \
  --model switchyard \
  --agent codex \
  --reasoning-effort xhigh \
  --n-tasks 1 \
  --n-concurrent 8 \
  --max-retries 2

```

## Locating Results and Logs

Upon completion, Switchyard writes artifacts to `benchmark/tb_runs/<run-id>/`. Query the final statistics to verify reproduction:

```bash
cat benchmark/tb_runs/<run-id>/routing_stats_final.json | jq .

```

This JSON contains per-tier call counts, token usage, and latency data cited in the Switchyard README.

## Summary

- Execute [`benchmark/prepare_harbor_dataset.py`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/prepare_harbor_dataset.py) to generate the closed-book Terminal-Bench 2.1 dataset.
- Establish upstream baselines using [`benchmark/run-baseline.sh`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/benchmark/run-baseline.sh) without a server configuration.
- Reproduce Switchyard numbers by targeting [`tb21-escalation-opus-glm-deepseek.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tb21-escalation-opus-glm-deepseek.toml) as the `--server-config` argument.
- Configure escalation thresholds via `confirmations`, `recent_turn_window`, and `window_message_chars` in the TOML.
- Inspect `benchmark/tb_runs/<run-id>/routing_stats_final.json` for token costs and accuracy metrics.

## Frequently Asked Questions

### What is the purpose of the escalation block in the TOML?

The escalation block in [`tb21-escalation-opus-glm-deepseek.toml`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/tb21-escalation-opus-glm-deepseek.toml) defines the policy for promoting requests from the weak tier (Kimi K2.7) to the strong tier (Claude Opus 4.7) when the classifier detects failure signals. It balances cost efficiency against accuracy by only invoking expensive models after confirmation thresholds are met.

### How does the `recent_turn_window` parameter affect routing decisions?

The `recent_turn_window` parameter specifies that Switchyard analyzes the last 28 conversation turns to determine if an escalation trigger is valid. This windowed context prevents premature escalation from transient errors and ensures the strong tier receives relevant diagnostic history capped at `window_message_chars`.

### Where are the benchmark artifacts stored after a run completes?

All run artifacts are persisted in `benchmark/tb_runs/<run-id>/` as implemented in the benchmark orchestration scripts. The [`routing_stats_final.json`](https://github.com/NVIDIA-NeMo/Switchyard/blob/main/routing_stats_final.json) file within this directory provides the structured data for accuracy and cost comparisons against the Terminal-Bench 2.1 baseline.

### Which models are assigned to the weak and strong tiers in this configuration?

According to the source TOML, the weak tier is assigned `moonshotai/kimi-k2.7-code`, the strong tier is assigned `anthropic/claude-opus-4.7`, and routing decisions are classified by `google/gemini-3.5-flash` as defined in the `[tiers]` and `[classifier]` sections.