DC-AE Tiling for 4K Resolution Inference with Limited GPU Memory: A Complete Guide
DC-AE tiling automatically splits high-resolution video frames into overlapping tiles during encoding and decoding, reducing peak GPU memory usage to single-tile levels and enabling 4K (3840×2160) inference on GPUs with as little as 8 GB of VRAM.
The DC-AE (dual-channel auto-encoder) in the NVlabs/Sana repository processes ultra-high-resolution content through an intelligent tiling system that requires no external preprocessing. By leveraging spatial and temporal tiling flags in the model configuration, users can run 4K resolution inference on consumer-grade GPUs without encountering out-of-memory errors.
How DC-AE Tiling Works
The tiling mechanism is fully encapsulated within the DCAE class and operates automatically during both encoding and decoding phases. The system processes tiles independently, ensuring that peak GPU memory never exceeds the requirements for a single patch rather than the full frame.
Configuration and Initialization
Tiling behavior is controlled through the DCVAEConfig dataclass, specifically within the video configuration block. According to the source code in diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py (lines 81-90), the key parameters include:
video.use_spatial_tiling– Enables 2D tiling across height and widthvideo.use_temporal_tiling– Enables tiling along the time dimension for video sequencesspatial_tile_sizeandtemporal_tile_size– Define the dimensions of each tile (default: 1024 pixels)tile_overlap_factor– Sets the blending overlap between tiles (default: 0.25 or 25%)
During model initialization (DCAE.__init__, lines 22-30), these configuration values are copied into instance attributes: self.use_spatial_tiling, self.use_temporal_tiling, and related size parameters.
The Encoding Pipeline
The DCAE.encode method (lines 40-50) dynamically selects the appropriate tiling strategy based on the configuration flags:
- Spatial tiling (
spatial_tiled_encode): Splits individual frames into 2D tiles (height × width) and processes each through the full encoder (self.encoder). This path handles 4K frames by processing 1024×1024 pixel regions sequentially. - Temporal tiling (
temporal_tiled_encode): Segments the video along the time axis, optionally nesting spatial tiling within each temporal chunk for long video sequences.
Each tile is encoded independently, producing latent representations that are later stitched together.
Seamless Blending
To eliminate visible seams between tiles, the implementation uses linear interpolation functions defined in lines 56-78 of dc_ae.py:
blend_v– Blends vertical overlapsblend_h– Blends horizontal overlapsblend_t– Blends temporal overlaps
These functions apply the tile_overlap_factor (default 25%) to create smooth transitions between adjacent tiles, ensuring the final latent map appears seamless despite being processed in chunks.
The Decoding Pipeline
The decoder mirrors the encoder's approach through spatial_tiled_decode and temporal_tiled_decode (lines 160-210). Each latent tile is decoded independently, and the resulting pixel-space tiles are blended using the same overlap functions before reconstruction. This symmetry ensures that memory usage remains constrained during both compression and decompression phases.
Enabling Tiling in Your Code
The model builder utility automatically invokes tiling when the configuration supports it. In diffusion/model/builder.py (line 352), the system calls:
ae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)
For manual instantiation, construct the VAE with a tiling-enabled configuration:
from diffusion.model.dc_ae.efficientvit.models.efficientvit.dc_ae import dc_vae_f32, DCAE
# Build configuration with tiling enabled
cfg = dc_vae_f32(
name="dc-vae-f32t4c128",
pretrained_path="hf://Efficient-Large-Model/SANA-Video_2B_480p/checkpoints/SANA_Video_2B_480p.pth",
) # Helper sets video.use_spatial_tiling=True and video.use_temporal_tiling=True
# Instantiate model
vae = DCAE(cfg)
# Optional: Customize tile dimensions for your GPU
vae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)
4K Resolution Inference Workflow
The following implementation demonstrates processing a 3840×2160 video while maintaining memory usage appropriate for 8-GB GPUs:
# 4k_tiling_demo.py
import torch
import imageio
from diffusion.model.dc_ae.efficientvit.models.efficientvit.dc_ae import dc_vae_f32, DCAE
# ----------------------------------------------------------------------
# 1. Load 4K video (3840×2160) as (T, C, H, W) tensor
# ----------------------------------------------------------------------
video_path = "sample_4k.mp4"
reader = imageio.get_reader(video_path)
frames = [
torch.from_numpy(frame).permute(2, 0, 1).float() / 255.0
for frame in reader
]
video = torch.stack(frames, dim=0) # Shape: (T, 3, 2160, 3840)
# ----------------------------------------------------------------------
# 2. Initialize VAE with tiling enabled
# ----------------------------------------------------------------------
cfg = dc_vae_f32(
name="dc-vae-f32t4c128",
pretrained_path="hf://Efficient-Large-Model/SANA-Video_2B_480p/checkpoints/SANA_Video_2B_480p.pth",
)
vae = DCAE(cfg)
# Force specific tile size (optional, default is 1024)
vae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)
# ----------------------------------------------------------------------
# 3. Encode with automatic tiling (memory-efficient)
# ----------------------------------------------------------------------
video = video.unsqueeze(0) # Add batch dimension: (1, T, 3, 2160, 3840)
latent = vae.encode(video) # Automatic spatial/temporal tiling applied
# ----------------------------------------------------------------------
# 4. Decode back to pixel space (also tiled)
# ----------------------------------------------------------------------
recon_video = vae.decode(latent)
print(f"Latent shape: {latent.shape}") # e.g., (1, 4, 40, 68, 120)
print(f"Reconstruction shape: {recon_video.shape}") # (1, T, 3, 2160, 3840)
Key implementation details:
- The full 4K frame never resides in GPU memory simultaneously; only individual 1024×1024 tiles are processed at any moment
- With the default compression ratio of 32, each spatial tile compresses to a 32×32 latent patch, further reducing memory overhead
- The pipeline in
app/sana_video_refiner_pipeline_diffusers.py(line 99) provides a production example of manually enabling tiling before inference
Memory Optimization Mechanism
The DC-AE tiling system achieves constant memory scaling regardless of input resolution. By processing tiles independently, peak GPU memory consumption is bounded by:
Memory_max = Memory_encoder(Tile_size) + Memory_overlap_buffer
For a 1024×1024 tile with 3 channels and float32 precision, this requires approximately 12 MB for the input tile, compared to 96 MB for a full 4K frame—a reduction that enables 4K video processing on hardware with limited VRAM. The latent representation receives proportionate savings, with each tile compressing to 32×32×4 channels at approximately 16 KB per tile.
Summary
- DC-AE tiling in NVlabs/Sana automatically splits high-resolution content into manageable tiles without manual preprocessing
- Configuration flags
video.use_spatial_tilingandvideo.use_temporal_tilinginDCVAEConfigcontrol the behavior - Memory usage scales with tile size (default 1024×1024) rather than frame resolution, enabling 4K inference on 8-GB GPUs
- Seamless reconstruction is achieved through linear blending functions (
blend_v,blend_h,blend_t) that handle 25% overlap regions - Implementation requires only calling
vae.enable_tiling()or using predefined configurations likedc_vae_f32()
Frequently Asked Questions
What is the default tile size for DC-AE tiling in Sana?
The default configuration uses 1024×1024 pixels for spatial tiles and 1024 frames for temporal segments. These values are set in the dc_vae_f32 configuration helper and can be overridden by passing tile_sample_min_height and tile_sample_min_width to the enable_tiling() method.
How does DC-AE prevent visible seams between tiles?
The system applies linear interpolation across overlapping regions using dedicated blending functions (blend_v, blend_h, and blend_t) defined in diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py. With a default tile_overlap_factor of 0.25 (25%), adjacent tiles blend smoothly to produce a seamless final output.
Can DC-AE tiling be used for still images as well as video?
Yes. While the configuration includes video.use_spatial_tiling and video.use_temporal_tiling flags, spatial tiling functions independently for single images (batched as (B, C, H, W)). Temporal tiling only activates when processing video tensors with a time dimension (B, T, C, H, W).
What is the minimum GPU memory required for 4K inference with DC-AE?
The system can process 4K (3840×2160) content on GPUs with 8 GB of VRAM by using the default 1024×1024 tile size. Memory usage remains constant regardless of frame dimensions because the encoder and decoder process only one tile at a time, with peak consumption determined solely by the tile dimensions and model width.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →