# How to Use RT-DETR with BoxMOT for Real-Time Detection and Tracking

> Learn how to use RT-DETR with BoxMOT for real-time detection and tracking. Integrate Hugging Face transformers for advanced object detection with ByteTrack or BoTSORT.

- Repository: [Mike/boxmot](https://github.com/mikel-brostrom/boxmot)
- Tags: how-to-guide
- Published: 2026-03-07

---

**BoxMOT supports RT-DETR (Real-Time Detection Transformer) through a dedicated `RTDetrStrategy` class that integrates Hugging Face transformers into the standard detector-plus-ReID pipeline, enabling transformer-based object detection with any BoxMOT tracker like ByteTrack or BoTSORT.**

The **mikel-brostrom/boxmot** repository provides a flexible multi-object tracking framework that accepts multiple detection backends. Using **RT-DETR with BoxMOT** allows you to leverage transformer-based architecture for real-time detection while maintaining compatibility with the library's tracking algorithms.

## Architecture Overview: How RT-DETR Integrates with BoxMOT

BoxMOT implements a **detector-plus-ReID pipeline** that abstracts away backend differences through strategy classes. RT-DETR integration follows a callback-driven pattern similar to YOLO-based models but uses Hugging Face transformers for the heavy lifting.

### Model Detection and Strategy Selection

When you specify a model path, `boxmot.detectors.get_yolo_inferer` inspects the string for RT-DETR identifiers. If the name contains any token from `RTDETR_MODELS` (such as `rtdetr_v2_r50vd`, `rtdetr_v2_r18vd`, or `rtdetr_v2_r101vd`), the function returns the `RTDetrStrategy` class.

In [`boxmot/detectors/__init__.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/detectors/__init__.py) (lines 26-34), the detection logic automatically routes transformer models to the appropriate handler:

```python

# From boxmot/detectors/__init__.py

RTDETR_MODELS = ['rtdetr_v2_r50vd', 'rtdetr_v2_r18vd', 'rtdetr_v2_r101vd']

def get_yolo_inferer(model_path):
    # Inspection logic that returns RTDetrStrategy for RT-DETR models

    if any(x in str(model_path) for x in RTDETR_MODELS):
        return RTDetrStrategy

```

### The RTDetrStrategy Implementation

The `RTDetrStrategy` class in [`boxmot/detectors/rtdetr.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/detectors/rtdetr.py) (lines 14-38) implements the minimal interface required by the pipeline: `__call__`, `preprocess`, `postprocess`, `warmup`, and path bookkeeping. It wraps the Hugging Face `RTDetrImageProcessor` and `RTDetrV2ForObjectDetection` objects.

During inference, the strategy's `__call__` method (lines 40-72) converts BGR `np.ndarray` inputs to PIL images, runs the transformer, and returns detections as a tensor of shape `[batch, N, 6]` where columns represent `[x1, y1, x2, y2, confidence, class]`.

### Pipeline Integration via Callbacks

The `DetectorReIDPipeline` in [`boxmot/engine/inference.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/engine/inference.py) (lines 86-99) registers internal callbacks for non-Ultralytics models:

- **`on_predict_start`** → `_setup_custom_model` loads the selected strategy (e.g., RT-DETR)
- **`on_predict_batch_start`** → `_update_custom_paths` stores source image filenames for result objects

This architecture allows RT-DETR to feed detections into BoxMOT's tracking layer (ByteTrack, BoTSORT, etc.) exactly as with YOLO models.

## Installation and Dependencies

RT-DETR requires the Hugging Face transformers library and timm. While BoxMOT installs these automatically on first use, pre-install them for faster startup:

```bash
pip install "transformers[torch]" timm

```

The dependencies provide:
- **`RTDetrImageProcessor`** – handles image preprocessing and tensor conversion
- **`RTDetrV2ForObjectDetection`** – the transformer model implementation

## Usage Examples

### Command-Line Interface (CLI)

Run real-time tracking with a webcam using RT-DETR as the detector:

```bash
uv run python -m boxmot.engine.cli track \
    --source 0 \
    --yolo-model rtdetr_v2_r50vd \
    --tracker bytetrack \
    --conf 0.4 \
    --show

```

BoxMOT downloads the pretrained model from the Hugging Face Hub (`PekingU/rtdetr_v2_r50vd`) automatically, then streams frames through the transformer backbone into the selected tracker.

### Python API Implementation

For programmatic control, initialize `DetectorReIDPipeline` with an RT-DETR model identifier:

```python
from pathlib import Path
from boxmot.engine.inference import DetectorReIDPipeline
import cv2

# Initialize pipeline with RT-DETR

pipeline = DetectorReIDPipeline(
    yolo_model_path="rtdetr_v2_r18vd",  # Triggers RTDetrStrategy

    device="cpu",                        # Use "cuda:0" for GPU

    half=False,
)

# Video capture loop

cap = cv2.VideoCapture(0)
while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break

    # Run detection + tracking

    results = pipeline(frame)  # Returns list of ultralytics.Results

    
    # Access detections: [N, 6] tensor → [x1, y1, x2, y2, conf, cls]

    for r in results:
        for *xyxy, conf, cls in r.boxes.cpu().numpy():
            cv2.rectangle(
                frame, 
                (int(xyxy[0]), int(xyxy[1])), 
                (int(xyxy[2]), int(xyxy[3])), 
                (0, 255, 0), 
                2
            )
    
    cv2.imshow("RT-DETR + BoxMOT", frame)
    if cv2.waitKey(1) & 0xFF == 27:  # ESC to quit

        break

cap.release()
cv2.destroyAllWindows()

```

The pipeline internally calls `RTDetrStrategy.__call__`, handles image-path bookkeeping via `_update_custom_paths`, and returns Ultralytics-compatible `Results` objects.

### Filtering Specific Classes

RT-DETR returns all 80 COCO classes by default. Filter for specific objects (e.g., persons and cars only) by setting the `classes` parameter:

```python
pipeline = DetectorReIDPipeline(
    yolo_model_path="rtdetr_v2_r101vd",
    device="cuda",
    half=True,
)

# Filter for person (0) and car (2) classes

pipeline.yolo.args.classes = [0, 2]

```

The `postprocess` method in [`boxmot/detectors/rtdetr.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/detectors/rtdetr.py) (lines 105-108) applies this filter using `torch.isin` before wrapping detections into `Results` objects.

## Key Implementation Details

- **Output Format**: RT-DETR returns tensors shaped `[batch, N, 6]` containing `[x1, y1, x2, y2, confidence, class]` coordinates in absolute pixel values
- **Results Wrapping**: The `postprocess` method (lines 93-112 in [`rtdetr.py`](https://github.com/mikel-brostrom/boxmot/blob/main/rtdetr.py)) converts raw tensors into Ultralytics `Results` objects, preserving the `names` dictionary from the model configuration for class label mapping
- **Warmup**: The `warmup` method prepares the model for inference batch sizes, ensuring consistent latency for real-time applications
- **Path Tracking**: The `_update_custom_paths` callback maintains the relationship between batch indices and source filenames, enabling proper metadata attachment to tracked objects

## Summary

- **BoxMOT** integrates RT-DETR through the `RTDetrStrategy` class, implementing the standard detector interface used by YOLO models.
- **Automatic detection** occurs in [`boxmot/detectors/__init__.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/detectors/__init__.py) when model names contain `rtdetr_v2_r50vd`, `rtdetr_v2_r18vd`, or `rtdetr_v2_r101vd`.
- **Dependencies** include `transformers[torch]` and `timm` for Hugging Face model support.
- **Inference flow** converts BGR numpy arrays to PIL, processes through `RTDetrV2ForObjectDetection`, and returns `[x1, y1, x2, y2, confidence, class]` tensors wrapped in Ultralytics `Results` objects.
- **Compatibility** extends to all BoxMOT trackers (ByteTrack, BoTSORT, etc.) without additional configuration.

## Frequently Asked Questions

### Can I use RT-DETR with any BoxMOT tracker?

Yes. Once the `DetectorReIDPipeline` processes RT-DETR outputs into standard `Results` objects, they feed into any tracking algorithm implemented in BoxMOT (ByteTrack, BoTSORT, DeepSORT, etc.) exactly like YOLO detections. The tracker receives bounding boxes, confidence scores, and class IDs in the expected format.

### What are the supported RT-DETR model variants?

BoxMOT recognizes three model identifiers: `rtdetr_v2_r50vd` (ResNet-50 backbone), `rtdetr_v2_r18vd` (ResNet-18 for faster inference), and `rtdetr_v2_r101vd` (ResNet-101 for higher accuracy). These correspond to pretrained models hosted on the Hugging Face Hub under the PekingU organization.

### Does RT-DETR require GPU acceleration?

No, though recommended for real-time performance. The `RTDetrStrategy` accepts a `device` parameter ("cpu", "cuda:0", etc.) and handles tensor placement automatically. CPU inference works for lower frame rates or high-resolution analysis, while GPU acceleration maintains real-time speeds with the transformer backbone.

### How does RT-DETR output differ from YOLO in BoxMOT?

Both backends return identical `Results` objects containing `boxes` tensors shaped `[N, 6]`. However, RT-DETR uses `RTDetrImageProcessor` for preprocessing (handling resizing and normalization specific to transformers) and returns confidence scores from the transformer decoder rather than YOLO's objectness-based predictions. The `postprocess` method in [`boxmot/detectors/rtdetr.py`](https://github.com/mikel-brostrom/boxmot/blob/main/boxmot/detectors/rtdetr.py) normalizes these differences before the tracking stage.