Parallel Box Decoding in LocateAnything: From Sequential Tokens to Atomic Box Prediction
Parallel Box Decoding (PBD) treats each bounding box as an atomic unit, predicting the full coordinate set (x₁, y₁, x₂, y₂) in a single forward pass rather than generating coordinates token-by-token, delivering 2×–6× speedup over autoregressive methods.
Parallel Box Decoding is the core architectural innovation powering LocateAnything, an open-source vision-language grounding model from the NVlabs/Eagle repository. Unlike traditional approaches that serialize bounding boxes into sequential coordinate tokens, PBD predicts complete box geometries in parallel. This shift from autoregressive generation to atomic unit prediction eliminates throughput bottlenecks while preserving geometric coherence.
The Problem with Token-by-Token Decoding
Traditional vision-language grounding models rely on Next Token Prediction (NTP) to decode bounding boxes. In this approach, the model serializes a box into a sequence of coordinate tokens—typically x₁ → y₁ → x₂ → y₂—and generates them sequentially.
This design creates two critical drawbacks according to the source code in Embodied/README.md:
-
Throughput bottleneck: Each coordinate requires a separate forward pass, severely limiting the number of boxes the model can output per second.
-
Geometric incoherence: Because coordinates are learned independently in separate generation steps, the model may produce inconsistent or irregular box structures where the spatial relationships between corners break down.
How Parallel Box Decoding Works
Parallel Box Decoding solves these limitations by fundamentally changing the decoding granularity. Instead of treating individual coordinates as tokens, PBD treats the entire bounding box (or a single point) as an atomic unit.
In Embodied/README.md#L28-L33, the implementation details show that PBD predicts the full coordinate set (x₁, y₁, x₂, y₂)—or (x, y) for point targets—in one forward pass. This design preserves intra-box geometry because the model learns the spatial relationships between corners simultaneously rather than sequentially.
The performance impact is substantial. Empirical results in the NVlabs/Eagle repository demonstrate that LocateAnything achieves a 2×–6× speedup over NTP-based methods, scaling from approximately 12 boxes per second (BPS) to roughly 25 BPS on dense scenes.
LocateAnything's Hybrid Inference Pipeline
LocateAnything integrates PBD into a flexible hybrid inference pipeline that balances speed and reliability. As documented in Embodied/README.md#L90-L94, the system supports three distinct generation modes:
-
Fast Mode (MTP): Runs Parallel Box Decoding for all boxes, delivering maximum throughput. This is the pure PBD path where every box is predicted atomically.
-
Slow Mode (NTP): Uses traditional autoregressive decoding as a fallback when parallel output is malformed or ambiguous. This provides a reliability backstop for complex edge cases.
-
Hybrid Mode (Default): Combines both approaches. The model first attempts PBD for fast inference, then automatically falls back to NTP for any problematic boxes. This ensures both speed and accuracy without manual intervention.
Implementing Parallel Box Decoding in Practice
The worker implementation in Embodied/locateanything_worker.py#L64-L68 exposes these modes through the generation_mode argument. When generation_mode="fast" or using the default hybrid setting, the underlying model invokes the PBD path, producing full-box predictions in a single step.
Here is how to use Parallel Box Decoding in your own code:
from PIL import Image
from locateanything_worker import LocateAnythingWorker
# Load the model (fast/parallel box decoding is the default)
worker = LocateAnythingWorker("nvidia/LocateAnything-3B")
# 1. Fast detection using PBD (default hybrid mode)
img = Image.open("street.jpg").convert("RGB")
result = worker.detect(img, ["person", "car", "bicycle"])
print("Raw answer:", result["answer"])
# Parse the atomic box tokens into pixel coordinates
w, h = img.size
boxes = LocateAnythingWorker.parse_boxes(result["answer"], w, h)
print("Parsed boxes:", boxes)
# 2. Force pure fast mode (explicit PBD) – no fallback to NTP
fast_result = worker.predict(
img,
"Locate all the instances that matches the following description: person</c>car.</c>",
generation_mode="fast", # forces PBD only
)
print("Fast-only answer:", fast_result["answer"])
# 3. Hybrid mode (fast with NTP fallback) – the default
hybrid_result = worker.predict(
img,
"Locate all the instances that matches the following description: person</c>car.</c>",
generation_mode="hybrid",
)
print("Hybrid answer:", hybrid_result["answer"])
The LocateAnythingWorker.parse_boxes() method handles the conversion from atomic predictions to pixel coordinates, abstracting away the low-level tensor manipulation while preserving the geometric integrity provided by PBD.
Summary
-
Parallel Box Decoding treats entire bounding boxes as atomic units rather than sequences of coordinate tokens, enabling single-pass prediction of
(x₁, y₁, x₂, y₂). -
Performance gains are substantial, with LocateAnything achieving 2×–6× speedup over autoregressive methods, scaling from 12 BPS to approximately 25 BPS on dense scenes.
-
Hybrid inference in
Embodied/locateanything_worker.pyprovides three modes—Fast (MTP), Slow (NTP), and Hybrid—allowing users to trade off between maximum throughput and fallback reliability. -
API access requires only setting the
generation_modeparameter to"fast"or"hybrid"when callingworker.predict().
Frequently Asked Questions
What is the difference between PBD and NTP in LocateAnything?
NTP (Next Token Prediction) generates bounding boxes by predicting coordinates one at a time in an autoregressive sequence (x₁, then y₁, then x₂, then y₂), requiring multiple forward passes per box. PBD (Parallel Box Decoding) predicts the complete coordinate tuple in a single forward pass, treating the box as an indivisible unit. According to Embodied/README.md, PBD eliminates the sequential latency and maintains better geometric consistency between corners.
How do I enable Parallel Box Decoding in my code?
Set the generation_mode parameter to "fast" when calling the predict() method on your LocateAnythingWorker instance. As shown in Embodied/locateanything_worker.py#L64-L68, this forces the model to use the PBD path exclusively. Alternatively, use "hybrid" (the default) to attempt PBD first and fall back to autoregressive decoding only for ambiguous predictions.
Why does LocateAnything use a hybrid mode instead of pure PBD?
While Parallel Box Decoding is significantly faster, certain edge cases with ambiguous visual features may produce malformed boxes. The Hybrid Mode documented in Embodied/README.md#L90-L94 automatically detects these failures and reverts to the more reliable NTP method for those specific instances. This ensures maximum throughput without sacrificing accuracy on difficult examples.
What performance improvement does Parallel Box Decoding provide?
Empirical benchmarks in the NVlabs/Eagle repository show that PBD delivers a 2×–6× speedup compared to token-by-token methods. Specifically, LocateAnything scales from approximately 12 boxes per second (BPS) using NTP to roughly 25 BPS using PBD on dense scenes, while simultaneously improving geometric coherence by learning spatial relationships jointly rather than independently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →