How the Harbor Framework Facilitates Evaluation and Benchmarking of Agent Performance

The Harbor framework provides a plug-and-play evaluation harness that runs Deep Agents inside isolated sandbox environments, automatically executes benchmark suites like Terminal Bench 2.0, and records execution traces in the ATIF format for deterministic scoring and analysis.

The langchain-ai/deepagents repository introduces Harbor as a specialized evaluation infrastructure designed to solve the reproducibility crisis in agent benchmarking. By combining isolated sandbox backends with standardized trajectory logging, Harbor transforms subjective agent outputs into measurable, comparable performance metrics. This framework enables researchers and engineers to validate agent behavior against deterministic test suites while maintaining full observability through LangSmith integration.

Core Architecture of the Harbor Evaluation Harness

Isolated Sandbox Backend (HarborSandbox)

The foundation of deterministic evaluation lies in libs/harbor/deepagents_harbor/backend.py, which implements the HarborSandbox class. This backend normalizes operating-system interactions by executing shell commands, file operations, and utilities like ls, grep, and glob inside containerized environments such as Docker, Modal, or Daytona.

Each operation returns strongly-typed responses including ExecuteResponse for command output and ReadResult for file system queries. The sandbox guarantees that every agent run encounters identical OS-level tool availability, eliminating environmental drift that typically corrupts benchmark comparisons.

Agent Integration Layer (DeepAgentsWrapper)

Located in libs/harbor/deepagents_harbor/deepagents_wrapper.py, the DeepAgentsWrapper class serves as the bridge between Deep Agents and the Harbor infrastructure. During initialization, the wrapper constructs a custom system prompt that injects the current directory listing and file system state directly into the agent's context window.

The wrapper exposes two primary asynchronous methods: setup() for environment initialization and run() for executing instructions against the sandbox. This architecture ensures that the LLM plans actions against the exact file system state present in the sandbox, creating a tight feedback loop between perception and action.

Standardized Trajectory Logging and Scoring

ATIF Format and Execution Traces

After an agent completes its task sequence, the wrapper invokes _save_trajectory() to serialize the LangChain message stream into the Agent Trajectory Interchange Format (ATIF). This JSON schema captures every step, tool call, observation, and token usage metric in a machine-readable structure.

Harbor consumes these ATIF files to compute the reward score—a normalized float between 0 and 1 representing the benchmark pass rate. This canonical representation also powers visualization tools in the Harbor UI, allowing developers to inspect exactly where an agent succeeded or failed during execution.

LangSmith Integration for Experiment Tracking

When the environment variable LANGSMITH_EXPERIMENT is configured, Harbor automatically wraps each run in a LangSmith trace block. The system records inputs, model metadata (including temperature and Harbor session ID), and outputs as structured feedback entries tagged with harbor_reward.

This integration enables correlation between qualitative trace analysis and quantitative performance scores. Researchers can filter runs by reward thresholds, compare prompt variations across experiments, and version control agent configurations within the LangSmith ecosystem.

Implementing Benchmark Evaluations

Running Terminal Bench 2.0 via CLI

Harbor ships with the Terminal Bench 2.0 dataset containing approximately 90 tasks designed to test file manipulation, code execution, and system navigation capabilities. The following command initiates a standardized evaluation loop:

uv run harbor run \
  --agent-import-path deepagents_harbor:DeepAgentsWrapper \
  --dataset terminal-bench@2.0 \
  -n 10 \
  --jobs-dir jobs/terminal-bench \
  --env docker

This invocation launches a Docker sandbox, executes ten tasks from the dataset, and automatically validates outputs against ground truth. Upon completion, Harbor prints the aggregate reward score (e.g., Reward: 0.78) and writes detailed trajectories to the specified jobs directory.

Programmatic Evaluation with Python

For custom evaluation pipelines, instantiate the wrapper directly to control model parameters and logging destinations:

from pathlib import Path
from deepagents_harbor.deepagents_wrapper import DeepAgentsWrapper
from harbor.environments.docker import DockerEnvironment

# Initialize logging directory for ATIF outputs

logs_dir = Path("./harbor_logs")
logs_dir.mkdir(exist_ok=True)

# Configure wrapper with deterministic parameters

wrapper = DeepAgentsWrapper(
    logs_dir=logs_dir,
    model_name="gpt-4o-mini",
    temperature=0.0,
    verbose=True,
    use_cli_agent=True,
)

# Provision sandbox environment

environment = DockerEnvironment(image="ubuntu:22.04")
await wrapper.setup(environment)

# Execute evaluation task

await wrapper.run(
    instruction="Create a Python script that prints the Fibonacci sequence up to 20.",
    environment=environment,
    context=None,
)

After execution, the trajectory.json file in logs_dir contains the complete ATIF record ready for Harbor's scoring engine.

Publishing Results to LangSmith

The helper script at libs/harbor/scripts/harbor_langsmith.py streamlines the feedback loop between Harbor and LangSmith:

python scripts/harbor_langsmith.py add-feedback \
  jobs/terminal-bench/2025-12-02__16-25-40 \
  --project-name deepagents-baseline-v1

This utility extracts per-task reward scores from the Harbor job directory and creates searchable LangSmith feedback entries, enabling statistical analysis across multiple experiment runs.

Summary

  • Deterministic Sandboxing: The HarborSandbox backend in backend.py ensures consistent OS-level tool availability across Docker, Modal, and Daytona environments.
  • Standardized Logging: ATIF trajectory files generated by DeepAgentsWrapper provide the canonical format for reward calculation and performance visualization.
  • Automated Scoring: Terminal Bench 2.0 integration delivers 0-1 reward scores based on test pass rates, stored as LangSmith feedback under the harbor_reward key.
  • Full Observability: Native LangSmith tracing captures model metadata, session IDs, and execution context for reproducible experiment analysis.

Frequently Asked Questions

What is the ATIF format in Harbor?

The Agent Trajectory Interchange Format (ATIF) is a JSON schema defined in the Harbor framework that serializes complete agent execution histories. It includes every tool call, observation, LLM response, and token usage metric. Harbor uses ATIF files to calculate reward scores and render execution visualizations in the UI.

How does Harbor ensure reproducible benchmark results?

Harbor achieves reproducibility through the HarborSandbox class, which isolates agents in containerized environments with deterministic file system states. By normalizing shell command execution and system tool availability across Docker, Modal, or Daytona backends, Harbor eliminates environmental variables that typically cause benchmark variance.

Can I use Harbor with custom datasets beyond Terminal Bench 2.0?

Yes. While Harbor ships with Terminal Bench 2.0 (--dataset terminal-bench@2.0), the framework supports custom benchmark suites. You can extend the DeepAgentsWrapper in deepagents_wrapper.py to handle different task formats, provided your dataset includes validation logic that returns a 0-1 reward score compatible with Harbor's feedback system.

What information does Harbor log to LangSmith?

Harbor records comprehensive metadata including the model name, temperature settings, Harbor session ID, experiment tags, and the final reward score (0-1). Each run creates a trace block containing inputs, outputs, and the harbor_reward feedback entry, enabling filtered queries and comparative analysis across agent configurations in the LangSmith interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →