# Eagle Tokenizer Configuration for Multimodal Inputs: A Complete Guide

> Discover the Eagle tokenizer configuration for multimodal inputs. Learn how Eagle uses Hugging Face AutoTokenizer with special image tokens and padding for vision-language encoding.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**Eagle uses a Hugging Face `AutoTokenizer` configured with `trust_remote_code=True`, augmented with image-specific special tokens (`<IMG_CONTEXT>`, `<IMG_START>`, `<IMG_END>`), forced left-side padding, and a reserved `image_token_index` to seamlessly encode vision-language inputs.**

The NVlabs/Eagle repository implements a sophisticated multimodal architecture that processes both textual and visual data through a unified token stream. Understanding the Eagle tokenizer configuration is essential for reproducing the model's vision-language capabilities or adapting the codebase for custom multimodal datasets.

## Base Tokenizer Initialization

Eagle initializes its tokenizer using the standard Hugging Face `AutoTokenizer` class with specific flags to support custom model attributes. In [`Embodied/locateanything_worker.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/locateanything_worker.py) at line 53, the tokenizer is loaded with `trust_remote_code=True` to expose the extra image-related tokens defined in the model repository.

```python
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained(
    model_path,
    trust_remote_code=True,
    add_eos_token=False,
    use_fast=False,
)

```

Setting `use_fast=False` ensures compatibility with the custom token manipulation logic implemented throughout the Eagle pipeline, particularly when handling special multimodal delimiters.

## Special Image Token Augmentation

The tokenizer configuration extends the base vocabulary with a fixed set of image-related special tokens. According to the source code in [`Embodied/eaglevl/utils/locany/processing_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/processing_locateanything.py) (lines 344-358), Eagle adds the following tokens to handle visual inputs:

- **`<IMG_CONTEXT>`** – The default placeholder token for each image patch
- **`<IMG_START>`** and **`<IMG_END>`** – Delimiters that mark the beginning and end of image token sequences
- **Structure tokens** – `</box>`, `</ref>`, `<|im_start|>`, and `오` (special Korean character) for grounding and referring tasks

These tokens are mapped to specific IDs and stored in the `LocateAnythingConfig` class. As defined in [`Embodied/eaglevl/utils/locany/configuration_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/configuration_locateanything.py) at line 58, the `image_token_index` (e.g., `151667`) provides a hardcoded reference for locating image placeholders within the token stream.

## Padding and Sequence Length Configuration

Eagle modifies two critical tokenizer behaviors to accommodate long multimodal sequences. First, the `padding_side` is explicitly set to `"left"` before any tokenization occurs, as seen in [`Embodied/evaluation/inference_sspro_ddp.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/evaluation/inference_sspro_ddp.py) at line 154. This left-padding configuration is required because image tokens are prepended to the input sequence.

Second, the tokenizer's maximum sequence length is overwritten to match training configurations. In [`Embodied/eaglevl/train/locany_finetune_magi_stream.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/train/locany_finetune_magi_stream.py) (lines 1305-1311), the code sets:

```python
tokenizer.model_max_length = training_args.max_seq_length  # Typically 2048 or 4096

```

This ensures the model can handle extended image-text sequences during both training and inference.

## Vocabulary Expansion and Embedding Resizing

During fine-tuning, Eagle dynamically expands the tokenizer vocabulary to support task-specific tokens. The [`locany_finetune_magi_stream.py`](https://github.com/NVlabs/Eagle/blob/main/locany_finetune_magi_stream.py) script (lines 1313-1326) adds additional special tokens including:

- `<box>` and `<ref>` – For bounding box and reference predictions
- `<null>` – Null token for missing annotations
- `<text_mask>` – For text region masking

After adding these tokens, the language model head must be resized to match the new vocabulary size:

```python
model.language_model.resize_token_embeddings(len(tokenizer))

```

This operation, found at lines 1422-1427 in the same file, ensures the embedding matrix aligns with the expanded tokenizer.

## Tokenizer-Processor Coupling

Eagle bundles the configured tokenizer with an image processor into a unified `LocateAnythingProcessor` class. As implemented in [`Embodied/eaglevl/utils/locany/processing_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/processing_locateanything.py) (lines 329-363), this processor exposes the tokenizer, image processor, and precomputed special token IDs to the rest of the pipeline.

The vision projector built in [`Eagle/eagle/model/multimodal_projector/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_projector/builder.py) consumes these tokenized inputs, using the `image_token_index` to identify which positions in the token stream correspond to visual features.

## Practical Implementation Example

To recreate Eagle's tokenizer configuration in your own environment:

```python
from transformers import AutoTokenizer
from Eagle.eaglevl.utils.locany.processing_locateanything import LocateAnythingProcessor

# Initialize base tokenizer

tokenizer = AutoTokenizer.from_pretrained(
    "path/to/eagle-model",
    trust_remote_code=True,
    use_fast=False,
    add_eos_token=False,
)

# Add multimodal special tokens

special_tokens = [
    "<IMG_CONTEXT>", "<IMG_START>", "<IMG_END>",
    "</box>", "</ref>", "<|im_start|>", "오",
    "<box>", "<ref>", "<null>", "<text_mask>"
]
tokenizer.add_tokens(special_tokens, special_tokens=True)

# Configure for multimodal inputs

tokenizer.padding_side = "left"
tokenizer.model_max_length = 4096

# Create processor bundle

image_token_index = tokenizer.convert_tokens_to_ids("<IMG_CONTEXT>")
processor = LocateAnythingProcessor(
    tokenizer=tokenizer,
    image_processor=clip_image_processor,
    image_token="<IMG_CONTEXT>",
    image_start_token="<IMG_START>",
    image_end_token="<IMG_END>",
    image_token_index=image_token_index,
)

```

## Summary

Eagle's multimodal tokenizer configuration relies on several key architectural decisions:

- **Hugging Face Base** – Uses `AutoTokenizer` with `trust_remote_code=True` and `use_fast=False` for custom token support
- **Visual Vocabulary** – Injects `<IMG_CONTEXT>`, `<IMG_START>`, and `<IMG_END>` tokens to mark image regions
- **Left Padding** – Forces `padding_side="left"` to properly align prepended image tokens
- **Dynamic Expansion** – Supports runtime vocabulary growth via `add_tokens()` with corresponding embedding resizing
- **Hardcoded Indices** – Maintains a fixed `image_token_index` (e.g., 151667) for efficient image token lookup

## Frequently Asked Questions

### Why does Eagle require left-side padding for multimodal inputs?

Eagle uses left-side padding because image tokens are prepended to the text sequence before processing. When `tokenizer.padding_side = "left"` is set (as seen in [`Embodied/evaluation/inference_sspro_ddp.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/evaluation/inference_sspro_ddp.py)), the padding tokens appear on the left side of the tensor, ensuring that the actual image and text tokens remain right-aligned and contiguous. This alignment is critical for the causal attention mechanism to correctly associate visual features with their corresponding text descriptions.

### What is the purpose of the `<IMG_CONTEXT>` token in Eagle's tokenizer configuration?

The `<IMG_CONTEXT>` token serves as the primary placeholder for image patch embeddings within the token stream. According to [`Embodied/eaglevl/utils/locany/configuration_locateanything.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/utils/locany/configuration_locateanything.py), this token maps to a specific `image_token_index` (e.g., 151667) that the vision projector uses to identify which positions in the input sequence contain visual information. During inference, these placeholder tokens are replaced with the actual projected image features from the multimodal projector.

### How does Eagle handle vocabulary expansion when adding special tokens?

Eagle handles vocabulary expansion through a two-step process defined in [`Embodied/eaglevl/train/locany_finetune_magi_stream.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/train/locany_finetune_magi_stream.py). First, new tokens like `<box>`, `<ref>`, and `<null>` are added to the tokenizer using `tokenizer.add_tokens()`. Second, the language model's embedding matrix is resized via `model.language_model.resize_token_embeddings(len(tokenizer))` to match the new vocabulary size. This ensures that the model can generate and process these task-specific tokens without dimension mismatches.

### Can I use a different base tokenizer with Eagle's multimodal configuration?

While Eagle's architecture is designed around specific Hugging Face transformer bases (typically Qwen-2 or LLaMA-style models), you can adapt the configuration to other compatible tokenizers provided they support the necessary customization hooks. The critical requirements are: support for `trust_remote_code=True`, the ability to add special tokens via `add_tokens()`, and sufficient vocabulary space to accommodate the additional image-related tokens (approximately 10-15 new tokens). You must also ensure the `image_token_index` is updated in the configuration to reflect the new token IDs.