How to Benchmark Different Trackers Using BoxMOT Evaluation Framework

BoxMOT provides a unified command-line interface and Python API that automates downloading MOT datasets, running multiple tracking algorithms, and computing HOTA, MOTA, and IDF1 metrics via the TrackEval suite.

The BoxMOT repository by mikel-brostrom offers a reproducible pipeline to benchmark different trackers using the BoxMOT evaluation framework on standard datasets like MOT17, MOT20, and VisDrone. This framework orchestrates detection generation, embedding extraction, tracking inference, and metric calculation in a single workflow.

Prerequisites and Dataset Setup

Before running benchmarks, the framework automatically handles dataset preparation through the eval_init function in boxmot/engine/evaluator.py (lines 60-78). This downloads the TrackEval toolkit and resolves dataset URLs from configs/datasets/*.yaml, ensuring the MOT challenge data is available locally and args.source points to the correct path.

Running the Benchmark from the Command Line

The boxmot eval sub-command, defined in boxmot/engine/cli.py, serves as the primary entry point for all benchmarking operations. It handles model path normalization, multi-model support, and dispatches to the core evaluation engine.

Benchmark a Single Tracker

To evaluate a specific tracker like DeepOCSORT on MOT17-mini:

boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method deepocsort \
    --batch-size 16 \
    --auto-batch \
    --resume

This command triggers three internal stages orchestrated by boxmot/engine/evaluator.py: run_generate_dets_embs creates cached detections and embeddings, run_generate_mot_results executes the tracker update loop via boxmot/engine/tracker.py (lines 84-115), and trackeval computes the final metrics.

Compare Multiple Trackers Side-by-Side

The CLI supports plural model options via plural_model_options in cli.py, allowing simultaneous benchmarking of multiple algorithms:

boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method botsort \
    --tracking-method strongsort \
    --batch-size 16 \
    --auto-batch \
    --resume

Each tracker runs sequentially, with results stored in distinct sub-folders under runs/mot/. Evaluation executes once per tracker, producing separate metric tables for direct comparison.

Using the Python API for Programmatic Benchmarking

For Jupyter notebooks or automated sweeps, import the evaluator directly and construct a SimpleNamespace to mirror CLI arguments:

from boxmot.engine.evaluator import main as run_eval
from types import SimpleNamespace
from pathlib import Path

args = SimpleNamespace(
    source="MOT17-mini/train",
    yolo_model=[Path("weights/yolov8n.pt")],
    reid_model=[Path("weights/osnet_x0_25_msmt17.pt")],
    tracking_method=["botsort", "deepocsort"],
    batch_size=16,
    auto_batch=True,
    resume=True,
    device="cuda",
    project=Path("runs"),
    name="benchmark_test"
)

run_eval(args)

This approach enables scripted parameter sweeps across different YOLO versions, ReID models, or tracker configurations without shell invocation.

Understanding the Evaluation Pipeline Internals

The framework follows a modular execution flow orchestrated by boxmot/engine/evaluator.py. First, create_tracker in boxmot/trackers/tracker_zoo.py (lines 27-95) instantiates the tracker class from TRACKER_MAPPING using YAML configs in configs/trackers/.

Then, run_generate_mot_results iterates over video sequences, loading pre-computed detections and embeddings before calling tracker.update(). Finally, the trackeval method (lines 144-168) constructs a command for the external TrackEval run_mot_challenge.py script, while parse_mot_results (lines 124-139) extracts HOTA, MOTA, and IDF1 scores from the stdout into a structured dictionary for the final report.

Adding Custom Trackers to the Benchmark

To benchmark a custom implementation, register it in boxmot/trackers/tracker_zoo.py:

TRACKER_MAPPING["mytracker"] = "boxmot.trackers.mytracker.mytracker.MyTracker"

Create a corresponding configuration file at configs/trackers/mytracker.yaml following the existing schema with tracker-specific hyper-parameters. Then benchmark it using the standard CLI:

boxmot eval \
    --source MOT17-mini/train \
    --yolo-model yolov8n.pt \
    --reid-model osnet_x0_25_msmt17.pt \
    --tracking-method mytracker \
    --batch-size 32

Summary

  • The boxmot eval CLI command automates the complete benchmarking pipeline from dataset download to metric reporting.
  • Multiple trackers can be compared in a single command by specifying --tracking-method multiple times, with results stored in separate experiment folders.
  • The Python API via boxmot.engine.evaluator.main enables programmatic experimentation and hyperparameter sweeps using SimpleNamespace.
  • Internally, create_tracker handles tracker instantiation while trackeval interfaces with the TrackEval suite to compute standard MOT metrics.
  • Custom trackers integrate seamlessly by updating TRACKER_MAPPING in tracker_zoo.py and adding YAML configuration files to configs/trackers/.

Frequently Asked Questions

What metrics does BoxMOT report during evaluation?

BoxMOT calculates standard MOT challenge metrics including HOTA, MOTA, IDF1, false positives, false negatives, and ID switches. The parse_mot_results function in boxmot/engine/evaluator.py extracts these values from the TrackEval output and presents them as both per-sequence tables and overall summaries via the shared logger.

Can I benchmark on custom datasets besides MOT17 or MOT20?

Yes. Add a new YAML configuration file in configs/datasets/ defining the download URLs, split information, and class filters. The eval_init function in evaluator.py automatically resolves and downloads the data, provided it follows the MOT challenge format with ground truth annotations in the expected directory structure.

How does BoxMOT handle pre-computed detections and embeddings?

The framework caches detections and ReID embeddings in the dets_n_embs/<benchmark>/ directory to avoid recomputation. The run_generate_dets_embs step checks for existing files before running YOLO inference, and the tracker loading mechanism in boxmot/engine/tracker.py reads these cached files during the tracking phase, significantly speeding up repeated benchmarks.

Is it possible to resume an interrupted benchmark?

Yes. Passing the --resume flag allows the pipeline to skip completed sequences and previously generated detection/embedding files. The tracker results are written incrementally to the experiment folder, and evaluation only processes sequences with available result files, making large-scale benchmarks resilient to interruptions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →