How DeepOCSort Handles Occlusions Using Appearance Cues

DeepOCSort mitigates occlusions by combining motion predictions with appearance embeddings during data association and preserving the last appearance feature for tracks that temporarily receive no detections.

DeepOCSort, implemented in the mikel-brostrom/boxmot repository, enhances multi-object tracking robustness by integrating ReID embeddings with the OCSort motion framework. This hybrid approach allows the tracker to maintain identity consistency even when objects are temporarily obscured or occluded, leveraging semantic appearance cues alongside geometric motion predictions.

Embedding Extraction and Association Strategy

ReID Feature Extraction

During frame processing, DeepOCSort extracts re-identification (ReID) features for every detection using self.model.get_features in boxmot/trackers/deepocsort/deepocsort.py (lines 44-50) unless the embedding_off parameter disables this functionality. These appearance vectors provide a semantic signature that complements geometric motion cues, enabling identity preservation when bounding boxes overlap or disappear.

Multi-Component Cost Matrix

The tracker constructs a detection-to-track cost matrix from three distinct components within the associate function located in boxmot/utils/association.py. First, IoU overlap is computed via self.asso_func. Second, velocity-direction consistency measures angle differences between predicted and observed motion vectors (lines 88-108). Third, appearance similarity is calculated as the dot-product stage1_emb_cost = dets_embs @ trk_embs.T (lines 94-99), creating a cosine similarity matrix between detection and track embeddings that captures visual identity regardless of spatial position.

Adaptive Weighting for Robust Matching

To prevent ambiguous matches from corrupting the association during partial occlusions, DeepOCSort employs an adaptive-weight (AW) module via compute_aw_max_metric in boxmot/utils/association.py (lines 21-28, 30-34). This module re-weights the embedding cost matrix to emphasize high-confidence appearance matches while suppressing uncertain similarities. By down-weighting ambiguous embedding comparisons, the system reduces false associations with similar-looking distractors when motion predictions become unreliable.

Two-Stage Association with OCR Rescue

The tracker implements a cascaded matching strategy to handle varying levels of occlusion severity. The first stage uses the full multi-component cost matrix including embeddings. For unmatched detections and tracks, a second-stage OCR-based rescue revisits associations using IoU and, if enabled, the remaining embedding cost (emb_cost_left) as implemented in boxmot/trackers/deepocsort/deepocsort.py (lines 119-134). This fallback mechanism provides recovery opportunities when motion cues alone are insufficient, effectively rescuing tracks during short-term occlusions where appearance remains consistent but motion predictions drift.

Frozen Tracks and Appearance Preservation

When a track receives no matching detection—indicating occlusion—the tracker calls self.active_tracks[m].update(None) and marks the track as frozen by setting self.frozen = True (lines 77-81). Crucially, the track's last embedding (self.emb) remains unchanged (lines 122-128), while the Kalman filter continues predicting motion via the update mechanism in boxmot/motion/kalman_filters/aabb/xysr_kf.py. This preservation allows immediate re-association when the object re-emerges, as the stored appearance feature is reused in subsequent association steps without requiring re-initialization.

import torch
import numpy as np
from pathlib import Path
from boxmot.trackers.deepocsort.deepocsort import DeepOcSort

# Initialise the tracker (provide a ReID model checkpoint)

tracker = DeepOcSort(
    reid_weights=Path("weights/reid_model.pt"),
    device=torch.device("cpu"),
    half=False,
    delta_t=3,               # motion window

    embedding_off=False,     # enable appearance cues

    aw_off=False,           # enable adaptive weighting

)

# Process a video frame-by-frame

for frame_id, (img, detections) in enumerate(video_frames):
    """
    detections: Nx5 array [[x1, y1, x2, y2, score], ...]
    img:       HxWx3 uint8 image
    """
    # Run the tracker – it extracts embeddings internally

    tracks = tracker.update(detections, img)

    # `tracks` is an Nx7 array: [x1, y1, x2, y2, id, conf, cls]

    for trk in tracks:
        x1, y1, x2, y2, obj_id = trk[:5].astype(int)
        cv2.rectangle(img, (x1, y1), (x2, y2), (0, 255, 0), 2)
        cv2.putText(img, f"ID {obj_id}", (x1, y1 - 5),
                    cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 2)

    # visualise or write `img` ...

When an object is occluded, tracker.update calls track.update(None). The Kalman filter predicts the next position while the previously stored embedding (track.emb) stays unchanged, allowing the object to be re-matched once it becomes visible again.

Summary

  • DeepOCSort extracts ReID embeddings for every detection using self.model.get_features in boxmot/trackers/deepocsort/deepocsort.py unless disabled by embedding_off
  • The association cost matrix combines IoU, velocity-direction consistency, and appearance similarity (dets_embs @ trk_embs.T) in boxmot/utils/association.py
  • Adaptive weighting via compute_aw_max_metric filters ambiguous embedding matches to prevent identity switches
  • A two-stage association process includes an OCR rescue fallback for occluded objects using emb_cost_left
  • Tracks frozen during occlusion preserve their last appearance vector (self.emb) for seamless re-identification when the object reappears

Frequently Asked Questions

How does DeepOCSort extract appearance features?

DeepOCSort utilizes a dedicated ReID model loaded via the reid_weights parameter. During each frame update, the tracker calls self.model.get_features in boxmot/trackers/deepocsort/deepocsort.py to generate normalized embedding vectors for every bounding box detection. These features are cached and used in the association cost matrix unless the embedding_off flag is set to True.

What happens to appearance embeddings when a track is occluded?

When occlusion occurs and no detection matches a track, DeepOCSort invokes track.update(None) and freezes the track by setting self.frozen = True while preserving the existing self.emb attribute unchanged. As implemented in boxmot/trackers/deepocsort/deepocsort.py (lines 122-128), this preservation ensures the track maintains its identity signature during temporary disappearance, enabling correct re-association when the object becomes visible again.

How does the adaptive weighting module improve occlusion handling?

The adaptive-weight module, implemented as compute_aw_max_metric in boxmot/utils/association.py, dynamically scales the appearance cost matrix based on match confidence. It amplifies costs for high-similarity pairs while attenuating ambiguous matches, preventing false associations with similar-looking objects during partial occlusions. This selective emphasis on strong appearance cues reduces identity switches when motion predictions become unreliable.

Can DeepOCSort operate without appearance cues?

Yes. Setting embedding_off=True during tracker initialization disables ReID feature extraction and removes appearance terms from the association cost matrix. In this mode, DeepOCSort relies solely on IoU and velocity-direction consistency for matching, though occlusion handling performance degrades significantly without the preserved appearance embeddings described in the standard configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →