# How LocateAnything Performs Zero-Shot Object Detection in the Wild: A Technical Deep Dive

> Unlock zero-shot object detection with LocateAnything. Discover how this unified vision-language model achieves precise detection in the wild without fine-tuning. Learn the technical details.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: deep-dive
- Published: 2026-06-28

---

**LocateAnything is a unified vision-language model that detects objects in arbitrary images or videos without task-specific fine-tuning by encoding visual inputs with a frozen Vision Transformer, projecting features into the language model's embedding space, and generating structured bounding-box tokens through multi-token prediction.**

The NVlabs/Eagle repository introduces LocateAnything (LA), a system that eliminates the need for supervised training on specific object categories. Instead of relying on traditional detector heads, the model treats object detection as a language generation task, allowing it to locate any visual concept described in natural language. This approach enables true zero-shot object detection in the wild, operating on everything from static images to video streams without architecture changes.

## Visual Encoding with MoonViT

At the foundation of LocateAnything's zero-shot capability is a frozen Vision Transformer called **MoonViT**. This encoder processes raw pixels or sampled video frames and returns dense visual embeddings that capture spatial and semantic information.

The vision encoder is instantiated within `LocateAnythingForConditionalGeneration.__init__` (lines [001‑011] of [`Embodied/eaglevl/utils/locany/modeling_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/modeling_locateanything.py)). During the forward pass, the `extract_feature` method (lines [002‑003]) transforms input images into a sequence of patch embeddings. These embeddings remain frozen during training, ensuring that the model leverages pre-trained visual representations while adapting only the language interface for detection tasks.

### Feature Extraction Pipeline

The MoonViT backbone, implemented in [`Embodied/eaglevl/model/moon_vit/modeling_vit.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/model/moon_vit/modeling_vit.py), outputs patch-level features that preserve spatial location information. When processing video, the encoder samples frames to create a temporal sequence of visual tokens. These tokens serve as the raw material for the detection process, containing all necessary visual information without manual annotation or bounding box supervision.

## Aligning Vision and Language

Raw visual embeddings exist in a different dimensional space than the language model's hidden states. LocateAnything bridges this gap through a learned projection layer that enables the LLM to interpret visual content.

### Projecting Visual Embeddings

The MLP `self.mlp1` (lines [036‑040] of [`modeling_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/modeling_locateanything.py)) maps the MoonViT hidden size (`vit_hidden_size`) to the LLM hidden size (`llm_hidden_size`). This projection translates visual features into the same semantic space as text tokens, allowing the language model to attend to visual information using its standard self-attention mechanisms.

### Token Replacement Logic

During the forward pass, the model replaces special image-token placeholders with the projected visual embeddings. In `LocateAnythingPreTrainedModel.forward` (lines [188‑238]), when images are present (`has_images` is True), the code builds a list of visual embeddings (`filtered_vit_embeds`), concatenates them, runs them through `self.mlp1`, and overwrites the corresponding positions in `input_embeds` (lines [190‑235]).

This substitution occurs at the placeholder token index (`self.image_token_index`), effectively inserting the visual information where the text prompt requested it (e.g., at `<image-1>` markers).

## Zero-Shot Detection Generation

LocateAnything generates bounding boxes not through regression heads, but by predicting discrete tokens from a predefined detection vocabulary. This design allows the model to output structured detection data using the same autoregressive generation process as text.

### Bounding Box Vocabulary

The model is trained to output specific tokens that encode spatial information:
- `<box>` – marks the start of a detection box
- Coordinate tokens (`coord_start_token_id` through `coord_end_token_id`) – represent normalized box coordinates
- `<none>` – indicates "no object" or empty predictions
- `<im_end>` – marks the end of image processing

Because these tokens are part of the model's standard vocabulary, no additional detection head or task-specific dataset is required. The LLM learns to associate visual features with spatial locations through the unified training objective.

### Multi-Token Prediction Modes

During generation, the model can operate in three distinct modes: fast MTP (multi-token prediction), hybrid, and slow AR (autoregressive). In the **fast** and **hybrid** modes, the model predicts multiple tokens simultaneously to accelerate inference.

The sampling routine `_sample_token_in_mtp` (lines [015‑030] of [`modeling_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/modeling_locateanything.py)) extracts predicted tokens, translates them into box tokens using `handle_pattern`, and determines whether a box is empty or valid. When running in **hybrid** mode, the system automatically switches back to autoregressive decoding upon encountering a `box_end_token_id` (lines [050‑056]), allowing error recovery while maintaining speed.

## End-to-End Inference Pipeline

The `LocateAnythingProcessor` orchestrates the entire zero-shot inference workflow, handling media loading and token preparation.

The processor reads images or video from multiple sources—local paths, URLs, LMDB records, or base-64 data—through `fetch_image` and `fetch_video` (lines [031‑056] of [`Embodied/eaglevl/utils/locany/processing_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/processing_locateanything.py)). It then parses placeholders in the user prompt (such as `<image-1>` or `<video-1>`) and substitutes them with visual tokens using `replace_media_placeholder` (lines [091‑128]).

The combined token stream, containing both text and projected visual embeddings, flows into the LLM, which generates a textual description interleaved with detection tokens (e.g., "<box> <coord> ..."). Because the detection vocabulary is intrinsic to the model's pretraining, the system locates objects zero-shot in any image or video without downstream task adaptation.

## Practical Implementation

The following code demonstrates how to run zero-shot object detection using the LocateAnything model from the NVlabs/Eagle repository:

```python
from transformers import AutoProcessor, AutoModelForCausalLM
from eagle.eagle.model.eagle_arch import LocateAnythingForConditionalGeneration

# Load the processor and model (both are HF-compatible)

processor = AutoProcessor.from_pretrained("NVlabs/locate-anything")
model = LocateAnythingForConditionalGeneration.from_pretrained("NVlabs/locate-anything")

# Prepare a prompt that asks for detection

prompt = "Identify all objects in this picture: <image-1>"

# The placeholder <image-1> will be replaced by the processor with visual tokens

# Pass the image (local path, URL, or LMDB entry) together with the prompt

# The processor automatically loads the image and builds the token stream

inputs = processor(
    text=prompt, 
    images="https://example.com/wild_scene.jpg", 
    return_tensors="pt"
)

# Generate a response that contains detection tokens

output = model.generate(
    input_ids=inputs["input_ids"],
    pixel_values=inputs["pixel_values"],
    image_grid_hws=inputs["image_grid_hws"],
    generation_mode="hybrid",      # fast-MTP with AR fallback

    max_new_tokens=256,
)

# Decode the generated string

# Bounding-box tokens appear as <box> ... </box>

detected_text = processor.tokenizer.decode(output[0], skip_special_tokens=False)
print(detected_text)

```

## Summary

- **LocateAnything** performs zero-shot object detection by treating bounding boxes as language tokens rather than regression targets, eliminating the need for task-specific fine-tuning.
- The **MoonViT** vision encoder (frozen during training) extracts dense visual features from images or videos, which are projected into the LLM's embedding space via `self.mlp1`.
- The **token replacement mechanism** in `LocateAnythingPreTrainedModel.forward` substitutes placeholder tokens with projected visual embeddings, allowing the language model to attend to visual content.
- Detection outputs use a specialized vocabulary including `<box>`, coordinate tokens, and `<none>`, generated through multi-token prediction (MTP) modes for efficiency.
- The **hybrid generation mode** combines fast multi-token prediction with autoregressive fallback when encountering `box_end_token_id`, balancing speed and accuracy.
- The **LocateAnythingProcessor** handles end-to-end orchestration, from loading media via `fetch_image`/`fetch_video` to replacing placeholders through `replace_media_placeholder`.

## Frequently Asked Questions

### How does LocateAnything differ from traditional object detectors like YOLO or DETR?

Traditional detectors rely on specialized regression heads and anchor boxes trained on specific object categories. LocateAnything eliminates these components entirely, instead using the LLM's generative capabilities to output bounding box tokens. According to the NVlabs/Eagle source code, the model generates detection results by predicting discrete tokens (`<box>`, coordinate tokens, `<none>`) from its standard vocabulary, allowing it to detect novel objects never seen during training.

### What is the purpose of the hybrid generation mode?

The hybrid generation mode optimizes the speed-accuracy trade-off during inference. As implemented in [`Embodied/eaglevl/utils/locany/modeling_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/modeling_locateanything.py) (lines [050‑056]), the model begins with fast multi-token prediction (MTP) to accelerate generation, but automatically switches to autoregressive (AR) decoding when it encounters a `box_end_token_id`. This allows the model to recover from potential errors in the fast prediction phase while maintaining higher throughput than pure AR generation.

### Can LocateAnything process video inputs for zero-shot detection?

Yes, the architecture handles video through the same unified pipeline. The processor's `fetch_video` method (lines [031‑056] of [`processing_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/processing_locateanything.py)) samples video frames, and the MoonViT encoder processes these as temporal sequences of visual tokens. The model can then generate bounding boxes for objects across video frames using the same `<video-1>` placeholder substitution and detection token generation used for static images.

### Why is the vision encoder frozen during training?

The MoonViT encoder remains frozen to preserve pre-trained visual representations while the model learns to align vision with language. This design choice, reflected in the `extract_feature` implementation (lines [002‑003] of [`modeling_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/modeling_locateanything.py)), ensures that the vision backbone maintains robust general-purpose features. Only the projection MLP (`self.mlp1`) and the LLM parameters are updated, allowing the model to adapt to detection tasks without overfitting to specific visual domains or object categories.