# Eagle 2.5 Long-Context Training Strategy: Processing 28K Tokens Efficiently

> Discover Eagle 2.5's efficient long-context training strategy. Learn how its custom attention kernel and token packing process 28K tokens for stable, high-performance results.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: deep-dive
- Published: 2026-06-28

---

**Eagle 2.5 leverages a custom MagiAttention kernel, token packing with sliding-window chunking, and mixed-precision training with gradient checkpointing to enable stable training on sequences up to 28,000 tokens.**

The NVlabs/Eagle repository implements Eagle 2.5, a vision-language model designed for long-context multimodal reasoning. According to the source code, its long-context training strategy specifically targets sequences of up to 28,000 tokens through a combination of memory-efficient attention mechanisms and a two-stage training pipeline.

## Token Packing and Streaming Architecture

The foundation of Eagle 2.5's long-context capability lies in how it structures multimodal inputs. Images, video frames, and text are tokenized and consolidated into a single contiguous stream.

### Sliding-Window Chunk Strategy

The packing logic follows a "sliding-window + chunk" scheme documented in [`Embodied/document/STREAMING_PACKING.md`](https://github.com/NVlabs/Eagle/blob/main/Embodied/document/STREAMING_PACKING.md). This approach keeps the most recent context active while discarding older chunks that are no longer needed, allowing a 28K-token sequence to fit within GPU memory constraints. The implementation packs tokens so that the model can process long videos and documents without truncation.

## MagiAttention for Extended Context Windows

Standard scaled dot-product attention (SDPA) fails beyond 16K tokens. Eagle 2.5 solves this by switching to a specialized kernel when processing long sequences.

### Block-Sparse Attention on Hopper GPUs

The training pipeline automatically enables **MagiAttention** when sequence lengths exceed 16K tokens, falling back to standard SDPA for shorter inputs. According to the [`Embodied/README.md`](https://github.com/NVlabs/Eagle/blob/main/Embodied/README.md), MagiAttention implements a **block-sparse** attention pattern that reduces memory bandwidth while preserving full-attention quality. This kernel is optimized for Hopper and Blackwell GPU architectures.

Configuration is gated by the `USE_MAGI_ATTENTION` flag found in `Eagle2_5/eaglevl/configs/*.py`:

```python

# In training configuration files

USE_MAGI_ATTENTION = True  # Enables block-sparse kernel for >16K tokens

```

## Memory Optimization Techniques

Processing 28K-token sequences requires aggressive memory management. Eagle 2.5 employs two complementary strategies to keep the GPU memory footprint under 20GB.

### Mixed-Precision Training

The training runs in **BF16/FP16** mixed precision, which halves the activation memory compared to full FP32 training. This setting is defined in [`Eagle2_5/eaglevl/configs/defaults.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/eaglevl/configs/defaults.py).

### Gradient Checkpointing

To further reduce memory usage, gradient checkpointing is applied specifically to the encoder layers. This trades computation for memory by recomputing activations during the backward pass rather than storing them, enabling the 28K-token streams to process on single GPUs or 8-GPU data-parallel jobs.

## Two-Stage Training Pipeline

Eagle 2.5's long-context training is split into distinct phases to prevent instability and ensure convergence.

### Stage 1: Short-Context Pre-training

First, the vision-language backbone is pre-trained on image-only data using shorter contexts. This stage establishes robust visual representations before introducing the complexity of long temporal sequences.

### Stage 2: Long-Context Fine-Tuning

The second stage finetunes the model using the full 28K-token streams on video-plus-image tasks. This stage leverages the packed token format and MagiAttention. The shell script [`train_stage2.sh`](https://github.com/NVlabs/Eagle/blob/main/train_stage2.sh) in the `Eagle2_5` folder illustrates this workflow, as detailed in [`Eagle2_5/document/3.training.md`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/document/3.training.md).

```bash

# Example Stage 2 launch configuration

bash Eagle2_5/train_stage2.sh \
  --sequence_length 28000 \
  --use_magi_attention \
  --gradient_checkpointing

```

## Evaluation on 28K-Token Benchmarks

The model's long-context capabilities are validated on tasks requiring extended reasoning, specifically **ChartQA** configured for 28K-token contexts and long-video understanding benchmarks. The main [`README.md`](https://github.com/NVlabs/Eagle/blob/main/README.md) benchmark table explicitly lists the 28K-token entry, confirming the model maintains performance across the full context window without degradation.

## Summary

- **Token Packing**: Implements sliding-window chunking via [`STREAMING_PACKING.md`](https://github.com/NVlabs/Eagle/blob/main/STREAMING_PACKING.md) to fit 28K tokens in memory.
- **MagiAttention**: Switches to block-sparse attention kernels for sequences >16K tokens, optimized for Hopper/Blackwell GPUs.
- **Memory Efficiency**: Combines BF16/FP16 mixed precision with gradient checkpointing to maintain ≤20GB memory usage.
- **Two-Stage Training**: Pre-trains on short contexts before fine-tuning on full 28K-token streams using [`train_stage2.sh`](https://github.com/NVlabs/Eagle/blob/main/train_stage2.sh).
- **Hardware**: Requires Hopper or Blackwell generation GPUs for optimal MagiAttention performance.

## Frequently Asked Questions

### What is the maximum context length Eagle 2.5 supports during training?

Eagle 2.5 supports training on sequences up to **28,000 tokens** through the MagiAttention kernel and token packing strategies documented in the NVlabs/Eagle repository.

### How does MagiAttention differ from standard SDPA?

MagiAttention implements a **block-sparse** attention pattern rather than full dense attention, significantly reducing memory bandwidth requirements while maintaining model quality. It activates automatically when `USE_MAGI_ATTENTION` is enabled and sequence lengths exceed 16K tokens.

### What hardware is required for 28K token training?

The implementation requires **Hopper or Blackwell GPUs** to run the MagiAttention kernel efficiently. With mixed-precision training and gradient checkpointing, the memory footprint remains under 20GB, allowing training on single GPUs or 8-GPU data-parallel configurations.

### Where is the token packing logic implemented?

The packing logic is documented in [`Embodied/document/STREAMING_PACKING.md`](https://github.com/NVlabs/Eagle/blob/main/Embodied/document/STREAMING_PACKING.md) and follows a sliding-window chunk strategy that discards older context while preserving recent tokens, enabling 28K-token sequences to fit within GPU memory constraints.