How to Debug Vision Tower Loading Issues in Eagle Training: A Complete Guide
To debug vision tower loading issues in Eagle training, verify the checkpoint path exists, inspect state dict key filtering for incompatible prefixes, ensure the EVA-ViT configuration matches the checkpoint dimensions, and run a dummy forward pass before starting training.
Training multimodal models with the NVlabs/Eagle repository requires properly loading the EVA-ViT vision tower. When the EVAVITVisionTower fails to initialize, the training script aborts with cryptic errors about missing weights or shape mismatches. Understanding how to debug vision tower loading issues in Eagle training ensures your EVA-02 backbone initializes correctly with pretrained weights.
Verify Configuration Arguments
Before the tower loads, Eagle expects three critical arguments in the args namespace. Missing or incorrect values cause immediate failures in EVAVITVisionTower.__init__.
vision_tower: The model name (e.g.,eva02-l-16) passed to the tower constructor.vision_tower_pretrained_from: Absolute or relative path to the.pthcheckpoint.freeze_vision: Boolean flag determining whether gradients flow through the backbone.
Add debugging prints before initialization to confirm values:
print("Vision tower:", args.vision_tower)
print("Pretrained checkpoint:", args.vision_tower_pretrained_from)
print("Freeze vision:", args.freeze_vision)
Confirm the Checkpoint File Exists
In Eagle/eagle/model/multimodal_encoder/vision_models/eva_vit.py, the load_model method checks file existence at lines 66–79. If os.path.exists returns False, the method emits a warning and sets self.is_loaded = True without actually loading weights, causing silent failures later.
if not os.path.exists(self.args.vision_tower_pretrained_from):
warnings.warn("The vision tower weights for EVA‑02 vision tower does not exists...")
self.is_loaded = True
return
Debug tip: Convert relative paths to absolute paths before launching training:
import os
checkpoint_path = os.path.abspath(args.vision_tower_pretrained_from)
assert os.path.exists(checkpoint_path), f"Checkpoint not found: {checkpoint_path}"
args.vision_tower_pretrained_from = checkpoint_path
Inspect State Dict Key Filtering
When the checkpoint exists, load_model (lines 90–103) strips specific prefixes and removes "rope" keys. The logic removes "backbone.net." for detection checkpoints or "visual." for CLIP checkpoints, then filters out any key containing "rope".
kw = ""
if "det" in self.args.vision_tower_pretrained_from.lower():
kw = "backbone.net."
elif "clip" in self.args.vision_tower_pretrained_from.lower():
kw = "visual."
if kw:
if kw in k and ("rope" not in k):
new_params[k.replace(kw, "")] = v
else:
if "rope" not in k:
new_params[k] = v
Debug tip: Print the first ten keys before and after filtering to verify essential layers like patch_embed.proj.weight are preserved:
print("Original keys:", list(pretrained_params.keys())[:10])
print("Filtered keys:", list(new_params.keys())[:10])
If patch_embed or blocks keys are missing, adjust the prefix removal logic or rename checkpoint keys to match the EVAViT architecture.
Handle Incompatible Keys Warnings
After filtering, load_state_dict runs with strict=False (lines 104–108). The code warns about any mismatched keys that do not contain "rope":
incompatiblekeys = self.vision_tower.load_state_dict(new_params, strict=False)
for k in incompatiblekeys[0]:
if "rope" not in k:
warnings.warn(f"Find incompatible keys {k} in state dict.")
Debug tip: Capture these warnings in a clean Python session. Common mismatches include:
- Missing
patch_embed.proj.biaswhen the checkpoint uses bias but the model does not. - Mismatched
hidden_dimbetween the checkpoint and the model configuration. - Residual
pos_embedkeys with different sequence lengths.
Validate Image Processor Configuration
The CLIPImageProcessor initializes with hard-coded mean and standard deviation values (lines 67–71). Mismatched preprocessing causes NaNs or feature distribution shifts during training.
self.image_processor = CLIPImageProcessor(
crop_size={"height": self.args.input_image_size, "width": self.args.input_image_size},
size={'shortest_edge': self.args.input_image_size},
image_mean=[0.48145466, 0.4578275, 0.40821073],
image_std=[0.26862954, 0.26130258, 0.27577711])
Debug tip: Verify your data loader uses identical normalization constants. Inconsistent image_mean or image_std values between preprocessing and the checkpoint's training regime produce drastically different feature magnitudes.
Verify Frozen Parameters
When args.freeze_vision is True, the tower calls requires_grad_(False) at lines 110–112:
if self.freeze_vision:
self.vision_tower.requires_grad_(False)
Debug tip: After initialization, confirm gradients are disabled to prevent GPU memory waste:
for name, param in tower.vision_tower.named_parameters():
assert not param.requires_grad, f"Parameter {name} should be frozen"
Calling model.train() later on the parent module can inadvertently re-enable gradients, so explicitly set tower.vision_tower.eval() if you freeze the vision backbone.
Run Forward Pass Sanity Check
Test the loaded tower with a dummy tensor before launching distributed training. The forward implementation in EVAViT.forward (lines 127–138) processes patch embeddings through transformer blocks.
import torch
dummy = torch.randn(1, 3, args.input_image_size, args.input_image_size)
out = tower.forward(dummy)
print(f"Output shape: {out.shape}") # Expected: (1, hidden_dim, H', W')
If the shape mismatches expectations, verify that args.input_image_size matches the img_size passed to build_eva_vit. The configuration dictionary returned by build_eva_vit (lines 155–165) contains critical fields like hidden_dim and num_patches that must align with the checkpoint.
Common Vision Tower Loading Errors
| Symptom | Root Cause | Solution |
|---|---|---|
FileNotFoundError or "does not exist" warning |
Incorrect vision_tower_pretrained_from path |
Use absolute paths or verify relative paths from the launch directory |
| "Find incompatible keys" warning | Checkpoint architecture differs from model (patch size, hidden dim) | Regenerate checkpoint with matching build_eva_vit config or manually remap keys |
RuntimeError: size mismatch during forward |
Input resolution mismatch | Ensure args.input_image_size equals the img_size used in build_eva_vit |
| NaN loss values | Image normalization mismatch | Align preprocessing mean/std with CLIPImageProcessor defaults |
| GPU OOM with frozen tower | Gradients re-enabled via model.train() |
Explicitly call tower.vision_tower.eval() after initialization |
Ready-to-Use Debug Script
Drop this script into your training workflow to verify the vision tower before the main loop starts:
import warnings
import torch
import os
from Eagle.eagle.model.multimodal_encoder.vision_models.eva_vit import (
EVAVITVisionTower, build_eva_vit
)
def debug_vision_tower(args):
# Verify checkpoint exists
ckpt_path = os.path.abspath(args.vision_tower_pretrained_from)
assert os.path.exists(ckpt_path), f"Checkpoint not found: {ckpt_path}"
# Initialize tower
tower = EVAVITVisionTower(
vision_tower=args.vision_tower,
args=args,
delay_load=False
)
print("=> Tower loaded:", tower.is_loaded)
print("=> Vision config:", tower.config)
# Verify parameter shapes
for name, param in tower.vision_tower.named_parameters():
print(f"{name:40} {list(param.shape)}")
# Test forward pass
dummy = torch.randn(1, 3, args.input_image_size, args.input_image_size)
with torch.no_grad():
out = tower.forward(dummy)
print("=> Forward output shape:", out.shape)
# Verify frozen status
if args.freeze_vision:
for p in tower.vision_tower.parameters():
assert not p.requires_grad, "Frozen parameters have requires_grad=True"
# Example usage
if __name__ == "__main__":
class Args:
vision_tower = "eva02-l-16"
vision_tower_pretrained_from = "/path/to/eva02_l_16.pth"
input_image_size = 224
freeze_vision = False
debug_vision_tower(Args())
Summary
Debugging vision tower loading issues in Eagle training requires systematic verification of the checkpoint pipeline:
- Verify paths: Ensure
vision_tower_pretrained_frompoints to an existing file using absolute paths. - Inspect keys: Check that state dict filtering removes prefixes correctly while preserving core
EVAViTparameters. - Match configs: Align
input_image_size,patch_size, andhidden_dimbetween the model architecture and checkpoint. - Validate preprocessing: Use the exact
CLIPImageProcessormean and standard deviation values. - Test before training: Run the debug script to confirm tensor shapes and frozen status before launching distributed training.
Frequently Asked Questions
Why does Eagle show "vision tower weights does not exist" when the file is present?
The load_model method in eva_vit.py uses os.path.exists which resolves paths relative to the current working directory. If you launch training from a different directory than where the path is defined, the check fails. Use os.path.abspath() to resolve the full path before assigning it to args.vision_tower_pretrained_from.
How do I fix incompatible keys warnings when loading EVA-ViT checkpoints?
Incompatible keys occur when the checkpoint contains layers not present in the model definition, or vice versa. Check the kw prefix logic in EVAVITVisionTower.load_model (lines 90–103). If your checkpoint uses a different prefix than "backbone.net." or "visual.", modify the filtering logic or manually rename keys in the checkpoint to match the EVAViT layer names (e.g., patch_embed, blocks, norm).
What image size should I use for the Eagle vision tower to avoid shape mismatches?
The image size must match the img_size parameter passed to build_eva_vit and the input_image_size argument used to initialize the image processor. Common values are 224, 336, or 448 depending on the EVA-02 variant. Mismatches cause VisionRotaryEmbeddingFast or PatchEmbed shape errors during the forward pass.
Why are gradients still flowing through my frozen vision tower?
Setting freeze_vision=True only calls requires_grad_(False) during initialization. If you later call model.train() on the parent model or the vision tower itself, PyTorch may reset gradient states. Explicitly call tower.vision_tower.eval() and tower.vision_tower.requires_grad_(False) after any model mode changes, or verify with assert not param.requires_grad before the training loop starts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →