How BoxMOT Trackers Handle Occluded Objects: OcSort vs DeepOcSort Implementation
BoxMOT trackers handle occluded objects through a two-stage association strategy that leverages low-confidence detections, Kalman filter predictions, and appearance embeddings to maintain track identities during temporary disappearances.
The BoxMOT library by mikel-brostrom provides production-ready multi-object tracking implementations that treat occlusion as a first-class scenario. Unlike simple trackers that immediately terminate tracks upon missing detections, BoxMOT trackers employ sophisticated mechanisms to "surf" through occlusions and recover object identities when they re-emerge. This article examines the specific implementation details in OcSort and DeepOcSort, revealing how these trackers maintain robust tracking through temporary object disappearance.
Two-Stage Association: The Core Occlusion Handling Strategy
BoxMOT implements occlusion handling directly inside the tracker update loops through a two-stage association architecture. This approach separates detections by confidence level and attempts recovery matching in a second pass.
Stage 1: High-Confidence Matching
In the first association stage, trackers match high-confidence detections (scores above det_thresh) to existing active tracks. This uses the primary cost function—typically IoU-based association or a combination of IoU, velocity, and appearance embeddings.
Stage 2: Low-Confidence Recovery
Detections with scores between min_conf and det_thresh are reserved for a second-stage association. These low-confidence detections often correspond to partially occluded objects. After the first matching round, any unmatched tracks are compared against these lower-quality detections, giving the tracker a chance to recover objects that became partially or fully occluded in previous frames.
OcSort: BYTE-Style Occlusion Recovery
The OcSort tracker (boxmot/trackers/ocsort/ocsort.py) implements explicit occlusion logic through its BYTE-style second-stage association. This tracker filters detections based on configurable confidence thresholds and performs sequential matching rounds.
Detection Splitting Logic
OcSort splits incoming detections into high and low confidence groups at lines 73-78:
inds_low = confs > self.min_conf
inds_high = confs < self.det_thresh
inds_second = np.logical_and(inds_low, inds_high)
This creates dets_second—an array of detections that are too weak for primary matching but strong enough to indicate potential object presence during occlusion.
BYTE Second-Stage Implementation
When use_byte=True and unmatched tracks exist, OcSort executes the second-stage association at lines 33-41:
if self.use_byte and len(dets_second) > 0 and unmatched_trks.shape[0] > 0:
iou_left = self.asso_func(dets_second, u_trks)
matched_indices = linear_assignment(-iou_left)
The linear_assignment function rematches low-confidence detections to previously unmatched tracks using the same IoU cost function. If the IoU exceeds the threshold, the track receives an update at lines 47-58 via self.active_tracks[trk_ind].update(...), effectively re-linking the occluded object to its historical track ID.
Kalman Prediction During Occlusion
Tracks that receive no detection in either stage are still kept alive. At lines 93-95, unmatched tracks advance via the Kalman filter:
for m in unmatched_trks:
self.active_tracks[m].update(None, None, None)
This prediction step allows the track to "surf" through short occlusions for a configurable max_age number of frames.
DeepOcSort: Appearance-Driven Occlusion Handling
DeepOcSort (boxmot/trackers/deepocsort/deepocsort.py) extends the two-stage strategy with camera-motion compensation (CMC) and appearance embeddings, providing robustness against both occlusion and camera movement.
Camera-Motion Compensation Preprocessing
Before any association occurs, DeepOcSort applies geometric transforms to align track predictions with the current camera pose at lines 52-57:
transform = self.cmc.apply(img, dets[:, :4])
for trk in self.active_tracks:
trk.apply_affine_correction(transform)
This correction ensures that Kalman-predicted bounding boxes remain spatially consistent even when the camera moves, reducing false mismatches during occlusion recovery.
Combined Cost Association
The first-stage association at lines 99-107 fuses multiple cues:
matched, unmatched_dets, unmatched_trks = associate(
...,
stage1_emb_cost,
self.w_association_emb,
...
)
Unlike OcSort, DeepOcSort incorporates appearance embeddings (stage1_emb_cost) alongside IoU and velocity, allowing the tracker to maintain identity even when spatial overlap is minimal due to occlusion.
Second-Stage IoU Recovery
After the first pass, remaining unmatched detections undergo a second matching round at lines 121-132:
iou_left = self.asso_func(left_dets, left_trks)
rematched_indices = linear_assignment(-iou_left)
This stage uses IoU-only matching to rescue tracks that might have been missed in the first stage due to embedding noise or motion blur during occlusion.
Kalman Filtering: The Foundation of Occlusion Survival
Both trackers rely on the KalmanBoxTracker implementation to maintain track states during occlusion. The Kalman filter predicts the next bounding box position even when no detection is available:
- In OcSort:
KalmanBoxTracker.predict()advances the state vector (lines 71-84 in the Kalman implementation) - In DeepOcSort: The same prediction mechanism allows tracks to survive for
max_ageframes without updates
This predictive capability is essential for handling occlusions lasting multiple frames, as the tracker maintains an estimated position until the object reappears.
Practical Implementation Examples
Using OcSort for Occlusion-Resilient Tracking
import cv2
import numpy as np
from boxmot.trackers.ocsort.ocsort import OcSort
# Initialize with low confidence threshold for second-stage matching
tracker = OcSort(min_conf=0.1, use_byte=True, det_thresh=0.3, max_age=30)
for frame_idx, frame in enumerate(video_frames):
# detections: (N,5) array [x1, y1, x2, y2, score]
detections = detector(frame)
tracks = tracker.update(detections, frame)
# tracks: (M,9) [x1, y1, x2, y2, id, conf, cls, det_ind]
for trk in tracks:
x1, y1, x2, y2, obj_id = trk[:5]
cv2.rectangle(frame, (int(x1), int(y1)), (int(x2), int(y2)), (0,255,0), 2)
cv2.putText(frame, f'ID {int(obj_id)}', (int(x1), int(y1)-5),
cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0,255,0), 2)
Key configuration: Setting min_conf=0.1 and use_byte=True enables the second-stage association that recovers occluded objects with detection scores between 0.1 and 0.3.
Using DeepOcSort with Appearance Embeddings
import cv2
import torch
from pathlib import Path
from boxmot.trackers.deepocsort.deepocsort import DeepOcSort
reid_path = Path('weights/osnet_x0_25_market1501.pt')
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
tracker = DeepOcSort(
reid_weights=reid_path,
device=device,
w_association_emb=0.5, # Weight for appearance cost
det_thresh=0.4,
max_age=30,
)
for frame in video_frames:
detections = detector(frame) # (N,5) array
embeddings = tracker.model.get_features(detections[:, :4], frame)
tracks = tracker.update(detections, frame, embs=embeddings)
for trk in tracks:
x1, y1, x2, y2, obj_id = trk[:5]
cv2.rectangle(frame, (int(x1), int(y1)), (int(x2), int(y2)), (255,0,0), 2)
Occlusion handling mechanism: The w_association_emb=0.5 parameter balances spatial and appearance cues, allowing the tracker to re-identify occluded objects based on visual similarity even when the predicted Kalman position drifts.
Summary
BoxMOT trackers handle occluded objects through a layered architectural approach:
- Two-stage association separates high-confidence detections from low-confidence candidates, enabling recovery of partially visible objects
- Kalman filter prediction maintains track states during detection gaps, allowing survival through short-term occlusions configurable via
max_age - BYTE mechanism (OcSort) explicitly rematches unmatched tracks with low-confidence detections using IoU cost matrices
- Camera-motion compensation (DeepOcSort) aligns predictions with frame-to-frame camera movement before association
- Appearance embeddings (DeepOcSort) provide additional identity cues during spatial ambiguity caused by occlusion
- Source implementations in
boxmot/trackers/ocsort/ocsort.pyandboxmot/trackers/deepocsort/deepocsort.pyexpose these mechanisms through theupdate()andassociate()functions
Frequently Asked Questions
What is the two-stage association strategy in BoxMOT?
The two-stage association strategy is a occlusion-handling technique where trackers first match high-confidence detections to tracks, then perform a second matching round using low-confidence detections (scores between min_conf and det_thresh) to recover tracks that were not paired initially. This approach is implemented in boxmot/trackers/ocsort/ocsort.py (lines 33-41) and boxmot/trackers/deepocsort/deepocsort.py (lines 121-132).
How does the Kalman filter help with occluded objects?
The Kalman filter predicts the next bounding box position using motion models even when no detection is assigned to a track. In BoxMOT, unmatched tracks call update(None) which triggers the Kalman predict() step, allowing the tracker to maintain estimated positions for up to max_age frames. This "surfing" capability prevents track termination during brief occlusions.
What is the difference between OcSort and DeepOcSort occlusion handling?
OcSort relies primarily on IoU-based BYTE matching for occlusion recovery, using a second-stage association with low-confidence detections. DeepOcSort adds appearance embeddings and camera-motion compensation (CMC) to the association cost, making it more robust against occlusions involving camera movement or similar-looking objects. DeepOcSort uses the w_association_emb parameter to control the balance between spatial and appearance cues.
How do I configure occlusion handling parameters?
Key parameters for occlusion handling include:
min_conf: Minimum detection confidence considered for second-stage matching (default varies by tracker)det_thresh: Threshold separating high and low confidence detectionsmax_age: Maximum frames a track survives without detection updatesuse_byte: Boolean flag enabling the second-stage association in OcSortw_association_emb: Weight for appearance cost in DeepOcSort (0.0 to 1.0)
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →