# How to Benchmark Different Trackers Using BoxMOT Evaluation Framework

> Benchmark trackers using the BoxMOT evaluation framework. Automate dataset downloads, run algorithms, and compute HOTA, MOTA, and IDF1 metrics effortlessly.

- Repository: [Mike/boxmot](https://github.com/mikel-brostrom/boxmot)
- Tags: how-to-guide
- Published: 2026-03-07

---

**BoxMOT provides a unified command-line interface and Python API that automates downloading MOT datasets, running multiple tracking algorithms, and computing HOTA, MOTA, and IDF1 metrics via the TrackEval suite.**

The BoxMOT repository by mikel-brostrom offers a reproducible pipeline to benchmark different trackers using the BoxMOT evaluation framework on standard datasets like MOT17, MOT20, and VisDrone. This framework orchestrates detection generation, embedding extraction, tracking inference, and metric calculation in a single workflow.

## Prerequisites and Dataset Setup

Before running benchmarks, the framework automatically handles dataset preparation through the `eval_init` function in [`boxmot/engine/evaluator.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/evaluator.py) (lines 60-78). This downloads the TrackEval toolkit and resolves dataset URLs from `configs/datasets/*.yaml`, ensuring the MOT challenge data is available locally and `args.source` points to the correct path.

## Running the Benchmark from the Command Line

The `boxmot eval` sub-command, defined in [`boxmot/engine/cli.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/cli.py), serves as the primary entry point for all benchmarking operations. It handles model path normalization, multi-model support, and dispatches to the core evaluation engine.

### Benchmark a Single Tracker

To evaluate a specific tracker like DeepOCSORT on MOT17-mini:

```bash
boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method deepocsort \
    --batch-size 16 \
    --auto-batch \
    --resume

```

This command triggers three internal stages orchestrated by [`boxmot/engine/evaluator.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/evaluator.py): `run_generate_dets_embs` creates cached detections and embeddings, `run_generate_mot_results` executes the tracker update loop via [`boxmot/engine/tracker.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/tracker.py) (lines 84-115), and `trackeval` computes the final metrics.

### Compare Multiple Trackers Side-by-Side

The CLI supports plural model options via `plural_model_options` in [`cli.py`](https://github.com/mikel-brostrom/boxmot/blob/main/cli.py), allowing simultaneous benchmarking of multiple algorithms:

```bash
boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method botsort \
    --tracking-method strongsort \
    --batch-size 16 \
    --auto-batch \
    --resume

```

Each tracker runs sequentially, with results stored in distinct sub-folders under `runs/mot/`. Evaluation executes once per tracker, producing separate metric tables for direct comparison.

## Using the Python API for Programmatic Benchmarking

For Jupyter notebooks or automated sweeps, import the evaluator directly and construct a `SimpleNamespace` to mirror CLI arguments:

```python
from boxmot.engine.evaluator import main as run_eval
from types import SimpleNamespace
from pathlib import Path

args = SimpleNamespace(
    source="MOT17-mini/train",
    yolo_model=[Path("weights/yolov8n.pt")],
    reid_model=[Path("weights/osnet_x0_25_msmt17.pt")],
    tracking_method=["botsort", "deepocsort"],
    batch_size=16,
    auto_batch=True,
    resume=True,
    device="cuda",
    project=Path("runs"),
    name="benchmark_test"
)

run_eval(args)

```

This approach enables scripted parameter sweeps across different YOLO versions, ReID models, or tracker configurations without shell invocation.

## Understanding the Evaluation Pipeline Internals

The framework follows a modular execution flow orchestrated by [`boxmot/engine/evaluator.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/evaluator.py). First, `create_tracker` in [`boxmot/trackers/tracker_zoo.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/trackers/tracker_zoo.py) (lines 27-95) instantiates the tracker class from `TRACKER_MAPPING` using YAML configs in `configs/trackers/`. 

Then, `run_generate_mot_results` iterates over video sequences, loading pre-computed detections and embeddings before calling `tracker.update()`. Finally, the `trackeval` method (lines 144-168) constructs a command for the external TrackEval [`run_mot_challenge.py`](https://github.com/mikel-brostrom/boxmot/blob/main/run_mot_challenge.py) script, while `parse_mot_results` (lines 124-139) extracts HOTA, MOTA, and IDF1 scores from the stdout into a structured dictionary for the final report.

## Adding Custom Trackers to the Benchmark

To benchmark a custom implementation, register it in [`boxmot/trackers/tracker_zoo.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/trackers/tracker_zoo.py):

```python
TRACKER_MAPPING["mytracker"] = "boxmot.trackers.mytracker.mytracker.MyTracker"

```

Create a corresponding configuration file at [`configs/trackers/mytracker.yaml`](https://github.com/mikel-brostrom/boxmot/blob/main/configs/trackers/mytracker.yaml) following the existing schema with tracker-specific hyper-parameters. Then benchmark it using the standard CLI:

```bash
boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method mytracker \
    --batch-size 32

```

## Summary

- The `boxmot eval` CLI command automates the complete benchmarking pipeline from dataset download to metric reporting.
- Multiple trackers can be compared in a single command by specifying `--tracking-method` multiple times, with results stored in separate experiment folders.
- The Python API via `boxmot.engine.evaluator.main` enables programmatic experimentation and hyperparameter sweeps using `SimpleNamespace`.
- Internally, `create_tracker` handles tracker instantiation while `trackeval` interfaces with the TrackEval suite to compute standard MOT metrics.
- Custom trackers integrate seamlessly by updating `TRACKER_MAPPING` in [`tracker_zoo.py`](https://github.com/mikel-brostrom/boxmot/blob/main/tracker_zoo.py) and adding YAML configuration files to `configs/trackers/`.

## Frequently Asked Questions

### What metrics does BoxMOT report during evaluation?

BoxMOT calculates standard MOT challenge metrics including **HOTA**, **MOTA**, **IDF1**, false positives, false negatives, and ID switches. The `parse_mot_results` function in [`boxmot/engine/evaluator.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/evaluator.py) extracts these values from the TrackEval output and presents them as both per-sequence tables and overall summaries via the shared logger.

### Can I benchmark on custom datasets besides MOT17 or MOT20?

Yes. Add a new YAML configuration file in `configs/datasets/` defining the download URLs, split information, and class filters. The `eval_init` function in [`evaluator.py`](https://github.com/mikel-brostrom/boxmot/blob/main/evaluator.py) automatically resolves and downloads the data, provided it follows the MOT challenge format with ground truth annotations in the expected directory structure.

### How does BoxMOT handle pre-computed detections and embeddings?

The framework caches detections and ReID embeddings in the `dets_n_embs/<benchmark>/` directory to avoid recomputation. The `run_generate_dets_embs` step checks for existing files before running YOLO inference, and the tracker loading mechanism in [`boxmot/engine/tracker.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/tracker.py) reads these cached files during the tracking phase, significantly speeding up repeated benchmarks.

### Is it possible to resume an interrupted benchmark?

Yes. Passing the `--resume` flag allows the pipeline to skip completed sequences and previously generated detection/embedding files. The tracker results are written incrementally to the experiment folder, and evaluation only processes sequences with available result files, making large-scale benchmarks resilient to interruptions.