How LocateAnything Performs Zero-Shot Object Detection in the Wild: A Technical Deep Dive

LocateAnything is a unified vision-language model that detects objects in arbitrary images or videos without task-specific fine-tuning by encoding visual inputs with a frozen Vision Transformer, projecting features into the language model's embedding space, and generating structured bounding-box tokens through multi-token prediction.

The NVlabs/Eagle repository introduces LocateAnything (LA), a system that eliminates the need for supervised training on specific object categories. Instead of relying on traditional detector heads, the model treats object detection as a language generation task, allowing it to locate any visual concept described in natural language. This approach enables true zero-shot object detection in the wild, operating on everything from static images to video streams without architecture changes.

Visual Encoding with MoonViT

At the foundation of LocateAnything's zero-shot capability is a frozen Vision Transformer called MoonViT. This encoder processes raw pixels or sampled video frames and returns dense visual embeddings that capture spatial and semantic information.

The vision encoder is instantiated within LocateAnythingForConditionalGeneration.__init__ (lines [001‑011] of Embodied/eaglevl/utils/locany/modeling_locateanything.py). During the forward pass, the extract_feature method (lines [002‑003]) transforms input images into a sequence of patch embeddings. These embeddings remain frozen during training, ensuring that the model leverages pre-trained visual representations while adapting only the language interface for detection tasks.

Feature Extraction Pipeline

The MoonViT backbone, implemented in Embodied/eaglevl/model/moon_vit/modeling_vit.py, outputs patch-level features that preserve spatial location information. When processing video, the encoder samples frames to create a temporal sequence of visual tokens. These tokens serve as the raw material for the detection process, containing all necessary visual information without manual annotation or bounding box supervision.

Aligning Vision and Language

Raw visual embeddings exist in a different dimensional space than the language model's hidden states. LocateAnything bridges this gap through a learned projection layer that enables the LLM to interpret visual content.

Projecting Visual Embeddings

The MLP self.mlp1 (lines [036‑040] of modeling_locateanything.py) maps the MoonViT hidden size (vit_hidden_size) to the LLM hidden size (llm_hidden_size). This projection translates visual features into the same semantic space as text tokens, allowing the language model to attend to visual information using its standard self-attention mechanisms.

Token Replacement Logic

During the forward pass, the model replaces special image-token placeholders with the projected visual embeddings. In LocateAnythingPreTrainedModel.forward (lines [188‑238]), when images are present (has_images is True), the code builds a list of visual embeddings (filtered_vit_embeds), concatenates them, runs them through self.mlp1, and overwrites the corresponding positions in input_embeds (lines [190‑235]).

This substitution occurs at the placeholder token index (self.image_token_index), effectively inserting the visual information where the text prompt requested it (e.g., at <image-1> markers).

Zero-Shot Detection Generation

LocateAnything generates bounding boxes not through regression heads, but by predicting discrete tokens from a predefined detection vocabulary. This design allows the model to output structured detection data using the same autoregressive generation process as text.

Bounding Box Vocabulary

The model is trained to output specific tokens that encode spatial information:

  • <box> – marks the start of a detection box
  • Coordinate tokens (coord_start_token_id through coord_end_token_id) – represent normalized box coordinates
  • <none> – indicates "no object" or empty predictions
  • <im_end> – marks the end of image processing

Because these tokens are part of the model's standard vocabulary, no additional detection head or task-specific dataset is required. The LLM learns to associate visual features with spatial locations through the unified training objective.

Multi-Token Prediction Modes

During generation, the model can operate in three distinct modes: fast MTP (multi-token prediction), hybrid, and slow AR (autoregressive). In the fast and hybrid modes, the model predicts multiple tokens simultaneously to accelerate inference.

The sampling routine _sample_token_in_mtp (lines [015‑030] of modeling_locateanything.py) extracts predicted tokens, translates them into box tokens using handle_pattern, and determines whether a box is empty or valid. When running in hybrid mode, the system automatically switches back to autoregressive decoding upon encountering a box_end_token_id (lines [050‑056]), allowing error recovery while maintaining speed.

End-to-End Inference Pipeline

The LocateAnythingProcessor orchestrates the entire zero-shot inference workflow, handling media loading and token preparation.

The processor reads images or video from multiple sources—local paths, URLs, LMDB records, or base-64 data—through fetch_image and fetch_video (lines [031‑056] of Embodied/eaglevl/utils/locany/processing_locateanything.py). It then parses placeholders in the user prompt (such as <image-1> or <video-1>) and substitutes them with visual tokens using replace_media_placeholder (lines [091‑128]).

The combined token stream, containing both text and projected visual embeddings, flows into the LLM, which generates a textual description interleaved with detection tokens (e.g., " ..."). Because the detection vocabulary is intrinsic to the model's pretraining, the system locates objects zero-shot in any image or video without downstream task adaptation.

Practical Implementation

The following code demonstrates how to run zero-shot object detection using the LocateAnything model from the NVlabs/Eagle repository:

from transformers import AutoProcessor, AutoModelForCausalLM
from eagle.eagle.model.eagle_arch import LocateAnythingForConditionalGeneration

# Load the processor and model (both are HF-compatible)

processor = AutoProcessor.from_pretrained("NVlabs/locate-anything")
model = LocateAnythingForConditionalGeneration.from_pretrained("NVlabs/locate-anything")

# Prepare a prompt that asks for detection

prompt = "Identify all objects in this picture: <image-1>"

# The placeholder <image-1> will be replaced by the processor with visual tokens

# Pass the image (local path, URL, or LMDB entry) together with the prompt

# The processor automatically loads the image and builds the token stream

inputs = processor(
    text=prompt, 
    images="https://example.com/wild_scene.jpg", 
    return_tensors="pt"
)

# Generate a response that contains detection tokens

output = model.generate(
    input_ids=inputs["input_ids"],
    pixel_values=inputs["pixel_values"],
    image_grid_hws=inputs["image_grid_hws"],
    generation_mode="hybrid",      # fast-MTP with AR fallback

    max_new_tokens=256,
)

# Decode the generated string

# Bounding-box tokens appear as <box> ... </box>

detected_text = processor.tokenizer.decode(output[0], skip_special_tokens=False)
print(detected_text)

Summary

  • LocateAnything performs zero-shot object detection by treating bounding boxes as language tokens rather than regression targets, eliminating the need for task-specific fine-tuning.
  • The MoonViT vision encoder (frozen during training) extracts dense visual features from images or videos, which are projected into the LLM's embedding space via self.mlp1.
  • The token replacement mechanism in LocateAnythingPreTrainedModel.forward substitutes placeholder tokens with projected visual embeddings, allowing the language model to attend to visual content.
  • Detection outputs use a specialized vocabulary including <box>, coordinate tokens, and <none>, generated through multi-token prediction (MTP) modes for efficiency.
  • The hybrid generation mode combines fast multi-token prediction with autoregressive fallback when encountering box_end_token_id, balancing speed and accuracy.
  • The LocateAnythingProcessor handles end-to-end orchestration, from loading media via fetch_image/fetch_video to replacing placeholders through replace_media_placeholder.

Frequently Asked Questions

How does LocateAnything differ from traditional object detectors like YOLO or DETR?

Traditional detectors rely on specialized regression heads and anchor boxes trained on specific object categories. LocateAnything eliminates these components entirely, instead using the LLM's generative capabilities to output bounding box tokens. According to the NVlabs/Eagle source code, the model generates detection results by predicting discrete tokens (<box>, coordinate tokens, <none>) from its standard vocabulary, allowing it to detect novel objects never seen during training.

What is the purpose of the hybrid generation mode?

The hybrid generation mode optimizes the speed-accuracy trade-off during inference. As implemented in Embodied/eaglevl/utils/locany/modeling_locateanything.py (lines [050‑056]), the model begins with fast multi-token prediction (MTP) to accelerate generation, but automatically switches to autoregressive (AR) decoding when it encounters a box_end_token_id. This allows the model to recover from potential errors in the fast prediction phase while maintaining higher throughput than pure AR generation.

Can LocateAnything process video inputs for zero-shot detection?

Yes, the architecture handles video through the same unified pipeline. The processor's fetch_video method (lines [031‑056] of processing_locateanything.py) samples video frames, and the MoonViT encoder processes these as temporal sequences of visual tokens. The model can then generate bounding boxes for objects across video frames using the same <video-1> placeholder substitution and detection token generation used for static images.

Why is the vision encoder frozen during training?

The MoonViT encoder remains frozen to preserve pre-trained visual representations while the model learns to align vision with language. This design choice, reflected in the extract_feature implementation (lines [002‑003] of modeling_locateanything.py), ensures that the vision backbone maintains robust general-purpose features. Only the projection MLP (self.mlp1) and the LLM parameters are updated, allowing the model to adapt to detection tasks without overfitting to specific visual domains or object categories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →