How to Use SFSORT for Lightweight CPU-Efficient Tracking in BoxMOT

SFSORT is a pure NumPy, CPU-optimized tracker in BoxMOT that uses a two-stage data association pipeline with dynamic threshold tuning to track objects without deep learning inference.

SFSORT (Simple, Fast, SORT) provides a lightweight alternative to deep-learning trackers when you need real-time multi-object tracking on CPU-only hardware. As implemented in the mikel-brostrom/boxmot repository, this tracker requires only bounding box coordinates, confidence scores, and class IDs—making it ideal for edge deployment and resource-constrained environments.

How SFSORT Achieves CPU Efficiency

Two-Stage Data Association Pipeline

SFSORT follows a strict two-step matching strategy defined in boxmot/trackers/sfsort/sfsort.py. First, high-confidence detections (above high_th) are matched against all active and lost tracks using a cost matrix that blends IoU, DIoU, and box-size similarity via the calculate_cost method (lines 43-99). Second, remaining intermediate detections (between low_th and high_th) are matched only to unmatched tracks using IoU-only association (iou_only=True). This hierarchical approach minimizes expensive matrix operations.

Dynamic Threshold Tuning

Instead of fixed thresholds, SFSORT optionally adapts to detection density through the _dynamic_thresholds method (lines 50-62). When dynamic_tuning=True, the tracker automatically adjusts high_th, low_th, and new_track_th based on the number of detections in the current frame. This eliminates manual tuning across different scenes without adding GPU overhead.

Pure NumPy Implementation

The tracker stores each object in a lightweight Track dataclass (lines 23-34) containing only essential fields: bbox, track_id, conf, cls, and a simple state machine (Active, Lost_Central, Lost_Marginal). All matrix operations use NumPy arrays rather than Python loops, and critically, no appearance embeddings are computed—the algorithm ignores the img parameter entirely, ensuring zero deep-learning inference cost inside the tracking loop.

Implementing SFSORT in Your Pipeline

Installation and Setup

Install BoxMOT via pip to access the SFSORT implementation:

pip install boxmot

The default configuration resides in boxmot/configs/trackers/sfsort.yaml, which the CLI loads when you specify --tracker sfsort.

Basic Usage Example

Integrate SFSORT into your detection loop by instantiating the SFSORT class and calling update() for each frame:

import numpy as np
from boxmot.trackers.sfsort.sfsort import SFSORT

# -------------------------------------------------

# 1️⃣ Initialise the tracker (CPU‑friendly defaults)

# -------------------------------------------------

tracker = SFSORT(
    high_th=0.6,           # high‑confidence threshold

    low_th=0.1,            # low‑confidence threshold

    new_track_th=0.7,      # when to start a new track

    match_th_first=0.67,   # first‑stage IoU match threshold

    match_th_second=0.3,   # second‑stage IoU match threshold

    dynamic_tuning=True,   # adapt thresholds on the fly

    marginal_timeout=30,   # frames before marginal track is dropped

    central_timeout=60,    # frames before central track is dropped

)

# -------------------------------------------------

# 2️⃣ Prepare a detection array for a single frame

#    Format: [x1, y1, x2, y2, confidence, class_id]

# -------------------------------------------------

detections = np.array([
    [100,  50, 200, 180, 0.92, 0],   # person

    [300, 120, 380, 210, 0.78, 2],   # car

    # … more detections …

], dtype=np.float32)

# Dummy image (required by the API but not used by SFSORT)

dummy_img = np.zeros((720, 1280, 3), dtype=np.uint8)

# -------------------------------------------------

# 3️⃣ Update the tracker – obtains a list of active tracks

# -------------------------------------------------

tracks = tracker.update(dets=detections, img=dummy_img)

# -------------------------------------------------

# 4️⃣ Inspect the result

#    Each row: [x1, y1, x2, y2, track_id, confidence, class_id, det_index]

# -------------------------------------------------

for t in tracks:
    x1, y1, x2, y2, tid, conf, cls, det_idx = t
    print(f"Track {int(tid)} (class {int(cls)}): "
          f"bbox=({x1:.1f},{y1:.1f},{x2:.1f},{y2:.1f}) "
          f"conf={conf:.2f}")

Configuration Parameters

Tune these key parameters in boxmot/trackers/sfsort/sfsort.py to balance precision and recall:

  • high_th: Detections above this score trigger guaranteed association or new track creation
  • low_th: Detections below this score are ignored entirely
  • match_th_first and match_th_second: IoU thresholds for stage-one and stage-two matching
  • central_timeout and marginal_timeout: Frames before purging lost tracks, determined by margin checks (lines 64-71)

Summary

  • SFSORT provides CPU-efficient multi-object tracking through pure NumPy operations in boxmot/trackers/sfsort/sfsort.py.
  • The two-stage association pipeline processes high-confidence detections first, then intermediate ones, minimizing computational waste.
  • Dynamic threshold tuning (_dynamic_thresholds) adapts to scene density without manual intervention or GPU resources.
  • The tracker requires only bounding boxes and confidence scores, making it compatible with any CPU-based detector exported to ONNX or similar formats.
  • State management uses lightweight dataclasses with configurable timeouts for central versus marginal lost tracks.

Frequently Asked Questions

Does SFSORT require GPU acceleration?

No. SFSORT is explicitly designed for CPU-only execution. According to the source code in boxmot/trackers/sfsort/sfsort.py, the tracker performs no deep learning inference, ignores the input image array, and relies entirely on NumPy matrix operations for IoU calculations and data association.

How does SFSORT handle occlusions without appearance features?

The tracker relies on spatial coherence through the calculate_cost function (lines 43-99), which combines IoU, DIoU, and box-size similarity. When objects become occluded, tracks enter either Lost_Central or Lost_Marginal states based on their last known position, remaining alive for central_timeout or marginal_timeout frames (lines 64-71) before deletion. This geometric approach works best for short occlusions in low-crowd scenarios.

What detection formats are compatible with SFSORT?

SFSORT accepts a NumPy array of shape (N, 6) where each row contains [x1, y1, x2, y2, confidence, class_id]. The output format extends this to (M, 8) with columns [x1, y1, x2, y2, track_id, confidence, class_id, det_index]. Any detector producing bounding boxes in this format—such as YOLOv5, YOLOv8, or ONNX-exported models—integrates seamlessly.

When should I enable dynamic threshold tuning?

Enable dynamic_tuning=True when your scene exhibits variable detection density, such as traffic cameras showing sparse rural roads transitioning to dense urban intersections. The _dynamic_thresholds method (lines 50-62) automatically adjusts confidence cutoffs based on the number of detections, preventing ID switches in dense crowds while maintaining sensitivity in sparse scenes. Disable it for fixed-camera setups with consistent object density to ensure deterministic behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →