# Parallel Box Decoding in LocateAnything: From Sequential Tokens to Atomic Box Prediction

> Discover Parallel Box Decoding in LocateAnything. This method predicts entire bounding boxes at once for 2x-6x faster decoding than token-by-token approaches. Learn more!

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: deep-dive
- Published: 2026-06-28

---

**Parallel Box Decoding (PBD) treats each bounding box as an atomic unit, predicting the full coordinate set `(x₁, y₁, x₂, y₂)` in a single forward pass rather than generating coordinates token-by-token, delivering 2×–6× speedup over autoregressive methods.**

Parallel Box Decoding is the core architectural innovation powering LocateAnything, an open-source vision-language grounding model from the NVlabs/Eagle repository. Unlike traditional approaches that serialize bounding boxes into sequential coordinate tokens, PBD predicts complete box geometries in parallel. This shift from autoregressive generation to atomic unit prediction eliminates throughput bottlenecks while preserving geometric coherence.

## The Problem with Token-by-Token Decoding

Traditional vision-language grounding models rely on **Next Token Prediction (NTP)** to decode bounding boxes. In this approach, the model serializes a box into a sequence of coordinate tokens—typically `x₁ → y₁ → x₂ → y₂`—and generates them sequentially.

This design creates two critical drawbacks according to the source code in [`Embodied/README.md`](https://github.com/NVlabs/Eagle/blob/main/Embodied/README.md):

1. **Throughput bottleneck**: Each coordinate requires a separate forward pass, severely limiting the number of boxes the model can output per second.

2. **Geometric incoherence**: Because coordinates are learned independently in separate generation steps, the model may produce inconsistent or irregular box structures where the spatial relationships between corners break down.

## How Parallel Box Decoding Works

Parallel Box Decoding solves these limitations by fundamentally changing the decoding granularity. Instead of treating individual coordinates as tokens, PBD treats the **entire bounding box** (or a single point) as an atomic unit.

In `Embodied/README.md#L28-L33`, the implementation details show that PBD predicts the full coordinate set `(x₁, y₁, x₂, y₂)`—or `(x, y)` for point targets—in one forward pass. This design preserves intra-box geometry because the model learns the spatial relationships between corners simultaneously rather than sequentially.

The performance impact is substantial. Empirical results in the NVlabs/Eagle repository demonstrate that LocateAnything achieves a **2×–6× speedup** over NTP-based methods, scaling from approximately 12 boxes per second (BPS) to roughly 25 BPS on dense scenes.

## LocateAnything's Hybrid Inference Pipeline

LocateAnything integrates PBD into a flexible **hybrid inference pipeline** that balances speed and reliability. As documented in `Embodied/README.md#L90-L94`, the system supports three distinct generation modes:

- **Fast Mode (MTP)**: Runs Parallel Box Decoding for all boxes, delivering maximum throughput. This is the pure PBD path where every box is predicted atomically.

- **Slow Mode (NTP)**: Uses traditional autoregressive decoding as a fallback when parallel output is malformed or ambiguous. This provides a reliability backstop for complex edge cases.

- **Hybrid Mode (Default)**: Combines both approaches. The model first attempts PBD for fast inference, then automatically falls back to NTP for any problematic boxes. This ensures both speed and accuracy without manual intervention.

## Implementing Parallel Box Decoding in Practice

The worker implementation in `Embodied/locateanything_worker.py#L64-L68` exposes these modes through the `generation_mode` argument. When `generation_mode="fast"` or using the default hybrid setting, the underlying model invokes the PBD path, producing full-box predictions in a single step.

Here is how to use Parallel Box Decoding in your own code:

```python
from PIL import Image
from locateanything_worker import LocateAnythingWorker

# Load the model (fast/parallel box decoding is the default)

worker = LocateAnythingWorker("nvidia/LocateAnything-3B")

# 1. Fast detection using PBD (default hybrid mode)

img = Image.open("street.jpg").convert("RGB")
result = worker.detect(img, ["person", "car", "bicycle"])
print("Raw answer:", result["answer"])

# Parse the atomic box tokens into pixel coordinates

w, h = img.size
boxes = LocateAnythingWorker.parse_boxes(result["answer"], w, h)
print("Parsed boxes:", boxes)

# 2. Force pure fast mode (explicit PBD) – no fallback to NTP

fast_result = worker.predict(
    img,
    "Locate all the instances that matches the following description: person</c>car.</c>",
    generation_mode="fast",        # forces PBD only

)
print("Fast-only answer:", fast_result["answer"])

# 3. Hybrid mode (fast with NTP fallback) – the default

hybrid_result = worker.predict(
    img,
    "Locate all the instances that matches the following description: person</c>car.</c>",
    generation_mode="hybrid",
)
print("Hybrid answer:", hybrid_result["answer"])

```

The `LocateAnythingWorker.parse_boxes()` method handles the conversion from atomic predictions to pixel coordinates, abstracting away the low-level tensor manipulation while preserving the geometric integrity provided by PBD.

## Summary

- **Parallel Box Decoding** treats entire bounding boxes as atomic units rather than sequences of coordinate tokens, enabling single-pass prediction of `(x₁, y₁, x₂, y₂)`.

- **Performance gains** are substantial, with LocateAnything achieving 2×–6× speedup over autoregressive methods, scaling from 12 BPS to approximately 25 BPS on dense scenes.

- **Hybrid inference** in [`Embodied/locateanything_worker.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/locateanything_worker.py) provides three modes—Fast (MTP), Slow (NTP), and Hybrid—allowing users to trade off between maximum throughput and fallback reliability.

- **API access** requires only setting the `generation_mode` parameter to `"fast"` or `"hybrid"` when calling `worker.predict()`.

## Frequently Asked Questions

### What is the difference between PBD and NTP in LocateAnything?

**NTP (Next Token Prediction)** generates bounding boxes by predicting coordinates one at a time in an autoregressive sequence (`x₁`, then `y₁`, then `x₂`, then `y₂`), requiring multiple forward passes per box. **PBD (Parallel Box Decoding)** predicts the complete coordinate tuple in a single forward pass, treating the box as an indivisible unit. According to [`Embodied/README.md`](https://github.com/NVlabs/Eagle/blob/main/Embodied/README.md), PBD eliminates the sequential latency and maintains better geometric consistency between corners.

### How do I enable Parallel Box Decoding in my code?

Set the `generation_mode` parameter to `"fast"` when calling the `predict()` method on your `LocateAnythingWorker` instance. As shown in `Embodied/locateanything_worker.py#L64-L68`, this forces the model to use the PBD path exclusively. Alternatively, use `"hybrid"` (the default) to attempt PBD first and fall back to autoregressive decoding only for ambiguous predictions.

### Why does LocateAnything use a hybrid mode instead of pure PBD?

While Parallel Box Decoding is significantly faster, certain edge cases with ambiguous visual features may produce malformed boxes. The **Hybrid Mode** documented in `Embodied/README.md#L90-L94` automatically detects these failures and reverts to the more reliable NTP method for those specific instances. This ensures maximum throughput without sacrificing accuracy on difficult examples.

### What performance improvement does Parallel Box Decoding provide?

Empirical benchmarks in the NVlabs/Eagle repository show that PBD delivers a **2×–6× speedup** compared to token-by-token methods. Specifically, LocateAnything scales from approximately 12 boxes per second (BPS) using NTP to roughly 25 BPS using PBD on dense scenes, while simultaneously improving geometric coherence by learning spatial relationships jointly rather than independently.