How to Use RT-DETR with BoxMOT for Real-Time Detection and Tracking

BoxMOT supports RT-DETR (Real-Time Detection Transformer) through a dedicated RTDetrStrategy class that integrates Hugging Face transformers into the standard detector-plus-ReID pipeline, enabling transformer-based object detection with any BoxMOT tracker like ByteTrack or BoTSORT.

The mikel-brostrom/boxmot repository provides a flexible multi-object tracking framework that accepts multiple detection backends. Using RT-DETR with BoxMOT allows you to leverage transformer-based architecture for real-time detection while maintaining compatibility with the library's tracking algorithms.

Architecture Overview: How RT-DETR Integrates with BoxMOT

BoxMOT implements a detector-plus-ReID pipeline that abstracts away backend differences through strategy classes. RT-DETR integration follows a callback-driven pattern similar to YOLO-based models but uses Hugging Face transformers for the heavy lifting.

Model Detection and Strategy Selection

When you specify a model path, boxmot.detectors.get_yolo_inferer inspects the string for RT-DETR identifiers. If the name contains any token from RTDETR_MODELS (such as rtdetr_v2_r50vd, rtdetr_v2_r18vd, or rtdetr_v2_r101vd), the function returns the RTDetrStrategy class.

In boxmot/detectors/__init__.py (lines 26-34), the detection logic automatically routes transformer models to the appropriate handler:


# From boxmot/detectors/__init__.py

RTDETR_MODELS = ['rtdetr_v2_r50vd', 'rtdetr_v2_r18vd', 'rtdetr_v2_r101vd']

def get_yolo_inferer(model_path):
    # Inspection logic that returns RTDetrStrategy for RT-DETR models

    if any(x in str(model_path) for x in RTDETR_MODELS):
        return RTDetrStrategy

The RTDetrStrategy Implementation

The RTDetrStrategy class in boxmot/detectors/rtdetr.py (lines 14-38) implements the minimal interface required by the pipeline: __call__, preprocess, postprocess, warmup, and path bookkeeping. It wraps the Hugging Face RTDetrImageProcessor and RTDetrV2ForObjectDetection objects.

During inference, the strategy's __call__ method (lines 40-72) converts BGR np.ndarray inputs to PIL images, runs the transformer, and returns detections as a tensor of shape [batch, N, 6] where columns represent [x1, y1, x2, y2, confidence, class].

Pipeline Integration via Callbacks

The DetectorReIDPipeline in boxmot/engine/inference.py (lines 86-99) registers internal callbacks for non-Ultralytics models:

  • on_predict_start → _setup_custom_model loads the selected strategy (e.g., RT-DETR)
  • on_predict_batch_start → _update_custom_paths stores source image filenames for result objects

This architecture allows RT-DETR to feed detections into BoxMOT's tracking layer (ByteTrack, BoTSORT, etc.) exactly as with YOLO models.

Installation and Dependencies

RT-DETR requires the Hugging Face transformers library and timm. While BoxMOT installs these automatically on first use, pre-install them for faster startup:

pip install "transformers[torch]" timm

The dependencies provide:

  • RTDetrImageProcessor – handles image preprocessing and tensor conversion
  • RTDetrV2ForObjectDetection – the transformer model implementation

Usage Examples

Command-Line Interface (CLI)

Run real-time tracking with a webcam using RT-DETR as the detector:

uv run python -m boxmot.engine.cli track \
    --source 0 \
    --yolo-model rtdetr_v2_r50vd \
    --tracker bytetrack \
    --conf 0.4 \
    --show

BoxMOT downloads the pretrained model from the Hugging Face Hub (PekingU/rtdetr_v2_r50vd) automatically, then streams frames through the transformer backbone into the selected tracker.

Python API Implementation

For programmatic control, initialize DetectorReIDPipeline with an RT-DETR model identifier:

from pathlib import Path
from boxmot.engine.inference import DetectorReIDPipeline
import cv2

# Initialize pipeline with RT-DETR

pipeline = DetectorReIDPipeline(
    yolo_model_path="rtdetr_v2_r18vd",  # Triggers RTDetrStrategy

    device="cpu",                        # Use "cuda:0" for GPU

    half=False,
)

# Video capture loop

cap = cv2.VideoCapture(0)
while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break

    # Run detection + tracking

    results = pipeline(frame)  # Returns list of ultralytics.Results

    
    # Access detections: [N, 6] tensor → [x1, y1, x2, y2, conf, cls]

    for r in results:
        for *xyxy, conf, cls in r.boxes.cpu().numpy():
            cv2.rectangle(
                frame, 
                (int(xyxy[0]), int(xyxy[1])), 
                (int(xyxy[2]), int(xyxy[3])), 
                (0, 255, 0), 
                2
            )
    
    cv2.imshow("RT-DETR + BoxMOT", frame)
    if cv2.waitKey(1) & 0xFF == 27:  # ESC to quit

        break

cap.release()
cv2.destroyAllWindows()

The pipeline internally calls RTDetrStrategy.__call__, handles image-path bookkeeping via _update_custom_paths, and returns Ultralytics-compatible Results objects.

Filtering Specific Classes

RT-DETR returns all 80 COCO classes by default. Filter for specific objects (e.g., persons and cars only) by setting the classes parameter:

pipeline = DetectorReIDPipeline(
    yolo_model_path="rtdetr_v2_r101vd",
    device="cuda",
    half=True,
)

# Filter for person (0) and car (2) classes

pipeline.yolo.args.classes = [0, 2]

The postprocess method in boxmot/detectors/rtdetr.py (lines 105-108) applies this filter using torch.isin before wrapping detections into Results objects.

Key Implementation Details

  • Output Format: RT-DETR returns tensors shaped [batch, N, 6] containing [x1, y1, x2, y2, confidence, class] coordinates in absolute pixel values
  • Results Wrapping: The postprocess method (lines 93-112 in rtdetr.py) converts raw tensors into Ultralytics Results objects, preserving the names dictionary from the model configuration for class label mapping
  • Warmup: The warmup method prepares the model for inference batch sizes, ensuring consistent latency for real-time applications
  • Path Tracking: The _update_custom_paths callback maintains the relationship between batch indices and source filenames, enabling proper metadata attachment to tracked objects

Summary

  • BoxMOT integrates RT-DETR through the RTDetrStrategy class, implementing the standard detector interface used by YOLO models.
  • Automatic detection occurs in boxmot/detectors/__init__.py when model names contain rtdetr_v2_r50vd, rtdetr_v2_r18vd, or rtdetr_v2_r101vd.
  • Dependencies include transformers[torch] and timm for Hugging Face model support.
  • Inference flow converts BGR numpy arrays to PIL, processes through RTDetrV2ForObjectDetection, and returns [x1, y1, x2, y2, confidence, class] tensors wrapped in Ultralytics Results objects.
  • Compatibility extends to all BoxMOT trackers (ByteTrack, BoTSORT, etc.) without additional configuration.

Frequently Asked Questions

Can I use RT-DETR with any BoxMOT tracker?

Yes. Once the DetectorReIDPipeline processes RT-DETR outputs into standard Results objects, they feed into any tracking algorithm implemented in BoxMOT (ByteTrack, BoTSORT, DeepSORT, etc.) exactly like YOLO detections. The tracker receives bounding boxes, confidence scores, and class IDs in the expected format.

What are the supported RT-DETR model variants?

BoxMOT recognizes three model identifiers: rtdetr_v2_r50vd (ResNet-50 backbone), rtdetr_v2_r18vd (ResNet-18 for faster inference), and rtdetr_v2_r101vd (ResNet-101 for higher accuracy). These correspond to pretrained models hosted on the Hugging Face Hub under the PekingU organization.

Does RT-DETR require GPU acceleration?

No, though recommended for real-time performance. The RTDetrStrategy accepts a device parameter ("cpu", "cuda:0", etc.) and handles tensor placement automatically. CPU inference works for lower frame rates or high-resolution analysis, while GPU acceleration maintains real-time speeds with the transformer backbone.

How does RT-DETR output differ from YOLO in BoxMOT?

Both backends return identical Results objects containing boxes tensors shaped [N, 6]. However, RT-DETR uses RTDetrImageProcessor for preprocessing (handling resizing and normalization specific to transformers) and returns confidence scores from the transformer decoder rather than YOLO's objectness-based predictions. The postprocess method in boxmot/detectors/rtdetr.py normalizes these differences before the tracking stage.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →