# How to Debug Vision Tower Loading Issues in Eagle Training: A Complete Guide

> Debug Eagle training vision tower loading issues. Verify paths inspect keys check config and run a dummy pass. Master NVlabs/Eagle troubleshooting for efficient AI development.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**To debug vision tower loading issues in Eagle training, verify the checkpoint path exists, inspect state dict key filtering for incompatible prefixes, ensure the EVA-ViT configuration matches the checkpoint dimensions, and run a dummy forward pass before starting training.**

Training multimodal models with the NVlabs/Eagle repository requires properly loading the EVA-ViT vision tower. When the `EVAVITVisionTower` fails to initialize, the training script aborts with cryptic errors about missing weights or shape mismatches. Understanding how to debug vision tower loading issues in Eagle training ensures your EVA-02 backbone initializes correctly with pretrained weights.

## Verify Configuration Arguments

Before the tower loads, Eagle expects three critical arguments in the `args` namespace. Missing or incorrect values cause immediate failures in `EVAVITVisionTower.__init__`.

- **`vision_tower`**: The model name (e.g., `eva02-l-16`) passed to the tower constructor.
- **`vision_tower_pretrained_from`**: Absolute or relative path to the `.pth` checkpoint.
- **`freeze_vision`**: Boolean flag determining whether gradients flow through the backbone.

Add debugging prints before initialization to confirm values:

```python
print("Vision tower:", args.vision_tower)
print("Pretrained checkpoint:", args.vision_tower_pretrained_from)
print("Freeze vision:", args.freeze_vision)

```

## Confirm the Checkpoint File Exists

In [`Eagle/eagle/model/multimodal_encoder/vision_models/eva_vit.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_encoder/vision_models/eva_vit.py), the `load_model` method checks file existence at lines 66–79. If `os.path.exists` returns `False`, the method emits a warning and sets `self.is_loaded = True` without actually loading weights, causing silent failures later.

```python
if not os.path.exists(self.args.vision_tower_pretrained_from):
    warnings.warn("The vision tower weights for EVA‑02 vision tower does not exists...")
    self.is_loaded = True
    return

```

**Debug tip**: Convert relative paths to absolute paths before launching training:

```python
import os
checkpoint_path = os.path.abspath(args.vision_tower_pretrained_from)
assert os.path.exists(checkpoint_path), f"Checkpoint not found: {checkpoint_path}"
args.vision_tower_pretrained_from = checkpoint_path

```

## Inspect State Dict Key Filtering

When the checkpoint exists, `load_model` (lines 90–103) strips specific prefixes and removes `"rope"` keys. The logic removes `"backbone.net."` for detection checkpoints or `"visual."` for CLIP checkpoints, then filters out any key containing `"rope"`.

```python
kw = ""
if "det" in self.args.vision_tower_pretrained_from.lower():
    kw = "backbone.net."
elif "clip" in self.args.vision_tower_pretrained_from.lower():
    kw = "visual."

if kw:
    if kw in k and ("rope" not in k):
        new_params[k.replace(kw, "")] = v
else:
    if "rope" not in k:
        new_params[k] = v

```

**Debug tip**: Print the first ten keys before and after filtering to verify essential layers like `patch_embed.proj.weight` are preserved:

```python
print("Original keys:", list(pretrained_params.keys())[:10])
print("Filtered keys:", list(new_params.keys())[:10])

```

If `patch_embed` or `blocks` keys are missing, adjust the prefix removal logic or rename checkpoint keys to match the `EVAViT` architecture.

## Handle Incompatible Keys Warnings

After filtering, `load_state_dict` runs with `strict=False` (lines 104–108). The code warns about any mismatched keys that do not contain `"rope"`:

```python
incompatiblekeys = self.vision_tower.load_state_dict(new_params, strict=False)
for k in incompatiblekeys[0]:
    if "rope" not in k:
        warnings.warn(f"Find incompatible keys {k} in state dict.")

```

**Debug tip**: Capture these warnings in a clean Python session. Common mismatches include:
- Missing `patch_embed.proj.bias` when the checkpoint uses bias but the model does not.
- Mismatched `hidden_dim` between the checkpoint and the model configuration.
- Residual `pos_embed` keys with different sequence lengths.

## Validate Image Processor Configuration

The `CLIPImageProcessor` initializes with hard-coded mean and standard deviation values (lines 67–71). Mismatched preprocessing causes NaNs or feature distribution shifts during training.

```python
self.image_processor = CLIPImageProcessor(
    crop_size={"height": self.args.input_image_size, "width": self.args.input_image_size},
    size={'shortest_edge': self.args.input_image_size},
    image_mean=[0.48145466, 0.4578275, 0.40821073],
    image_std=[0.26862954, 0.26130258, 0.27577711])

```

**Debug tip**: Verify your data loader uses identical normalization constants. Inconsistent `image_mean` or `image_std` values between preprocessing and the checkpoint's training regime produce drastically different feature magnitudes.

## Verify Frozen Parameters

When `args.freeze_vision` is `True`, the tower calls `requires_grad_(False)` at lines 110–112:

```python
if self.freeze_vision:
    self.vision_tower.requires_grad_(False)

```

**Debug tip**: After initialization, confirm gradients are disabled to prevent GPU memory waste:

```python
for name, param in tower.vision_tower.named_parameters():
    assert not param.requires_grad, f"Parameter {name} should be frozen"

```

Calling `model.train()` later on the parent module can inadvertently re-enable gradients, so explicitly set `tower.vision_tower.eval()` if you freeze the vision backbone.

## Run Forward Pass Sanity Check

Test the loaded tower with a dummy tensor before launching distributed training. The forward implementation in `EVAViT.forward` (lines 127–138) processes patch embeddings through transformer blocks.

```python
import torch

dummy = torch.randn(1, 3, args.input_image_size, args.input_image_size)
out = tower.forward(dummy)
print(f"Output shape: {out.shape}")  # Expected: (1, hidden_dim, H', W')

```

If the shape mismatches expectations, verify that `args.input_image_size` matches the `img_size` passed to `build_eva_vit`. The configuration dictionary returned by `build_eva_vit` (lines 155–165) contains critical fields like `hidden_dim` and `num_patches` that must align with the checkpoint.

## Common Vision Tower Loading Errors

| Symptom | Root Cause | Solution |
|---------|------------|----------|
| `FileNotFoundError` or "does not exist" warning | Incorrect `vision_tower_pretrained_from` path | Use absolute paths or verify relative paths from the launch directory |
| "Find incompatible keys" warning | Checkpoint architecture differs from model (patch size, hidden dim) | Regenerate checkpoint with matching `build_eva_vit` config or manually remap keys |
| `RuntimeError: size mismatch` during forward | Input resolution mismatch | Ensure `args.input_image_size` equals the `img_size` used in `build_eva_vit` |
| NaN loss values | Image normalization mismatch | Align preprocessing mean/std with `CLIPImageProcessor` defaults |
| GPU OOM with frozen tower | Gradients re-enabled via `model.train()` | Explicitly call `tower.vision_tower.eval()` after initialization |

## Ready-to-Use Debug Script

Drop this script into your training workflow to verify the vision tower before the main loop starts:

```python
import warnings
import torch
import os
from Eagle.eagle.model.multimodal_encoder.vision_models.eva_vit import (
    EVAVITVisionTower, build_eva_vit
)

def debug_vision_tower(args):
    # Verify checkpoint exists

    ckpt_path = os.path.abspath(args.vision_tower_pretrained_from)
    assert os.path.exists(ckpt_path), f"Checkpoint not found: {ckpt_path}"
    
    # Initialize tower

    tower = EVAVITVisionTower(
        vision_tower=args.vision_tower,
        args=args,
        delay_load=False
    )
    
    print("=> Tower loaded:", tower.is_loaded)
    print("=> Vision config:", tower.config)
    
    # Verify parameter shapes

    for name, param in tower.vision_tower.named_parameters():
        print(f"{name:40} {list(param.shape)}")
    
    # Test forward pass

    dummy = torch.randn(1, 3, args.input_image_size, args.input_image_size)
    with torch.no_grad():
        out = tower.forward(dummy)
    print("=> Forward output shape:", out.shape)
    
    # Verify frozen status

    if args.freeze_vision:
        for p in tower.vision_tower.parameters():
            assert not p.requires_grad, "Frozen parameters have requires_grad=True"

# Example usage

if __name__ == "__main__":
    class Args:
        vision_tower = "eva02-l-16"
        vision_tower_pretrained_from = "/path/to/eva02_l_16.pth"
        input_image_size = 224
        freeze_vision = False
    
    debug_vision_tower(Args())

```

## Summary

Debugging vision tower loading issues in Eagle training requires systematic verification of the checkpoint pipeline:

- **Verify paths**: Ensure `vision_tower_pretrained_from` points to an existing file using absolute paths.
- **Inspect keys**: Check that state dict filtering removes prefixes correctly while preserving core `EVAViT` parameters.
- **Match configs**: Align `input_image_size`, `patch_size`, and `hidden_dim` between the model architecture and checkpoint.
- **Validate preprocessing**: Use the exact `CLIPImageProcessor` mean and standard deviation values.
- **Test before training**: Run the debug script to confirm tensor shapes and frozen status before launching distributed training.

## Frequently Asked Questions

### Why does Eagle show "vision tower weights does not exist" when the file is present?

The `load_model` method in [`eva_vit.py`](https://github.com/NVlabs/Eagle/blob/main/eva_vit.py) uses `os.path.exists` which resolves paths relative to the current working directory. If you launch training from a different directory than where the path is defined, the check fails. Use `os.path.abspath()` to resolve the full path before assigning it to `args.vision_tower_pretrained_from`.

### How do I fix incompatible keys warnings when loading EVA-ViT checkpoints?

Incompatible keys occur when the checkpoint contains layers not present in the model definition, or vice versa. Check the `kw` prefix logic in `EVAVITVisionTower.load_model` (lines 90–103). If your checkpoint uses a different prefix than `"backbone.net."` or `"visual."`, modify the filtering logic or manually rename keys in the checkpoint to match the `EVAViT` layer names (e.g., `patch_embed`, `blocks`, `norm`).

### What image size should I use for the Eagle vision tower to avoid shape mismatches?

The image size must match the `img_size` parameter passed to `build_eva_vit` and the `input_image_size` argument used to initialize the image processor. Common values are 224, 336, or 448 depending on the EVA-02 variant. Mismatches cause `VisionRotaryEmbeddingFast` or `PatchEmbed` shape errors during the forward pass.

### Why are gradients still flowing through my frozen vision tower?

Setting `freeze_vision=True` only calls `requires_grad_(False)` during initialization. If you later call `model.train()` on the parent model or the vision tower itself, PyTorch may reset gradient states. Explicitly call `tower.vision_tower.eval()` and `tower.vision_tower.requires_grad_(False)` after any model mode changes, or verify with `assert not param.requires_grad` before the training loop starts.