Data Post-Training Strategies in Eagle 2 for Frontier Vision-Language Models

Eagle 2.5 employs Progressive Mixed Post-Training and Information-First Sampling to scale context windows from 32K to 128K tokens while preserving visual fidelity in high-resolution images and long video sequences.

The NVlabs/Eagle repository introduces Eagle 2 (released as the Eagle 2.5 family), a frontier vision-language model architecture that leverages sophisticated data post-training strategies to handle extensive multimodal contexts. These techniques enable the model to process up to 512 video frames and 4K resolution images without sacrificing detail, addressing the critical challenge of maintaining information density across variable-length inputs.

Progressive Mixed Post-Training

Eagle 2.5 implements a Progressive Mixed Post-Training schedule that gradually expands the model's maximum context length. According to the Eagle 2.5 documentation in Eagle2_5/README.md (lines 33-35 and 66), the training process begins with shorter contexts and systematically increases the maximum token capacity from 32K up to 128K tokens.

This progressive approach allows the model to adapt to increasingly large multimodal windows during the post-training phase. By gradually exposing the model to longer sequences, the strategy improves robustness when handling variable-length inputs and maintains training stability across the extended context range.

Training Script Configuration

The progressive context expansion is configured through training script arguments. In scripts/train_stage2.sh, the max_position_embeddings parameter defines the context window sizes used during different stages of the post-training pipeline. This file serves as the entry point for initiating the progressive schedule that prepares the model for long-context inference.

Information-First Sampling Pipeline

The Information-First Sampling strategy ensures that visual detail is preserved even when token budgets are constrained. This pipeline consists of two complementary techniques implemented during the data preparation phase of post-training.

Image Area Preservation (IAP)

Image Area Preservation maintains the original image area and aspect ratio during preprocessing. As documented in Eagle2_5/README.md (lines 30-33), IAP ensures that high-resolution visual information remains intact regardless of the final token allocation. This technique prevents the loss of fine-grained details that typically occurs during aggressive resizing operations.

Automatic Degrade Sampling (ADS)

Following IAP, Automatic Degrade Sampling balances the remaining visual content against textual requirements. ADS ensures that full text captions are retained even when the total token budget is limited, dynamically adjusting the compression ratio to preserve critical information from both modalities. This dual-stage approach allows Eagle 2 to handle inputs up to 4K resolution while maintaining coherent text generation.

Loading Long-Context Models in Practice

After completing the post-training regimen, models can be loaded using the extended context configuration. The load_pretrained_model function in Eagle/model/builder.py automatically reads the model's configured maximum context length (e.g., 128K tokens after progressive post-training).

The following example demonstrates how to load an Eagle 2.5 checkpoint and process long-video inputs:

from eagle.model.builder import load_pretrained_model
from eagle.mm_utils import tokenizer_image_token, process_images
from eagle.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX

# Load a checkpoint that has finished progressive post-training

model_path = "NVEagle/Eagle-2.5-8B"
model_name = "Qwen2.5-8B-Instruct"
tokenizer, model, image_processor, context_len = load_pretrained_model(
    model_path, None, model_name, flash_attn=False, deepspeed=False
)

# Prepare a long-video prompt (example uses 256 frames)

prompt = "<VideoFrames256>" + DEFAULT_IMAGE_TOKEN + " Describe the scene."

# Convert prompt to token IDs using the tokenizer from Eagle/mm_utils.py

input_ids = tokenizer_image_token(
    prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt"
)

# Run generation with the expanded context window

output_ids = model.generate(
    input_ids.unsqueeze(0),
    max_new_tokens=256,
    do_sample=True,
    temperature=0.2,
    top_p=0.5,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

The tokenizer_image_token utility in Eagle/mm_utils.py handles the special image token insertion required for multimodal processing. For complete inference examples, refer to Eagle2_5/document/5.inference.md.

Summary

  • Progressive Mixed Post-Training gradually expands context windows from 32K to 128K tokens, enabling Eagle 2.5 to process long video sequences and extensive documents.
  • Information-First Sampling combines Image Area Preservation and Automatic Degrade Sampling to maintain visual detail in high-resolution inputs up to 4K.
  • Training configuration occurs in scripts/train_stage2.sh through the max_position_embeddings parameter, which controls the progressive context schedule.
  • The load_pretrained_model function in Eagle/model/builder.py automatically configures the inference pipeline to utilize the extended context capabilities achieved through post-training.

Frequently Asked Questions

What is progressive mixed post-training in Eagle 2?

Progressive mixed post-training is a training schedule that systematically increases the model's maximum context length from 32K tokens to 128K tokens. This gradual expansion, documented in Eagle2_5/README.md, allows the model to adapt to longer multimodal sequences without destabilizing the training process, ultimately enabling processing of up to 512 video frames.

How does Information-First Sampling preserve visual detail?

Information-First Sampling utilizes Image Area Preservation (IAP) to maintain original image dimensions and aspect ratios, followed by Automatic Degrade Sampling (ADS) to intelligently balance visual and textual token allocation. This ensures that high-resolution visual information is retained even when the total available context is constrained.

What is the maximum context length supported by Eagle 2.5?

Eagle 2.5 supports context windows up to 128K tokens after completing the progressive post-training regimen. This extends the base model's capacity from an initial 32K tokens, allowing the model to handle long-form video analysis and high-resolution image understanding simultaneously.

Where can I find the training scripts for Eagle 2 post-training?

The training entry point for configuring post-training parameters is located at scripts/train_stage2.sh in the NVlabs/Eagle repository. This script contains the max_position_embeddings arguments necessary to implement the progressive context expansion strategy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →