# DC-AE Tiling for 4K Resolution Inference with Limited GPU Memory: A Complete Guide

> Learn DC-AE tiling to perform 4K resolution inference on GPUs with only 8GB VRAM. This guide explains how tiling drastically reduces peak GPU memory usage.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: how-to-guide
- Published: 2026-05-19

---

**DC-AE tiling automatically splits high-resolution video frames into overlapping tiles during encoding and decoding, reducing peak GPU memory usage to single-tile levels and enabling 4K (3840×2160) inference on GPUs with as little as 8 GB of VRAM.**

The **DC-AE** (dual-channel auto-encoder) in the NVlabs/Sana repository processes ultra-high-resolution content through an intelligent tiling system that requires no external preprocessing. By leveraging spatial and temporal tiling flags in the model configuration, users can run 4K resolution inference on consumer-grade GPUs without encountering out-of-memory errors.

## How DC-AE Tiling Works

The tiling mechanism is fully encapsulated within the `DCAE` class and operates automatically during both encoding and decoding phases. The system processes tiles independently, ensuring that peak GPU memory never exceeds the requirements for a single patch rather than the full frame.

### Configuration and Initialization

Tiling behavior is controlled through the `DCVAEConfig` dataclass, specifically within the `video` configuration block. According to the source code in [`diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py) (lines 81-90), the key parameters include:

- `video.use_spatial_tiling` – Enables 2D tiling across height and width
- `video.use_temporal_tiling` – Enables tiling along the time dimension for video sequences
- `spatial_tile_size` and `temporal_tile_size` – Define the dimensions of each tile (default: 1024 pixels)
- `tile_overlap_factor` – Sets the blending overlap between tiles (default: 0.25 or 25%)

During model initialization (`DCAE.__init__`, lines 22-30), these configuration values are copied into instance attributes: `self.use_spatial_tiling`, `self.use_temporal_tiling`, and related size parameters.

### The Encoding Pipeline

The `DCAE.encode` method (lines 40-50) dynamically selects the appropriate tiling strategy based on the configuration flags:

- **Spatial tiling** (`spatial_tiled_encode`): Splits individual frames into 2D tiles (height × width) and processes each through the full encoder (`self.encoder`). This path handles 4K frames by processing 1024×1024 pixel regions sequentially.
- **Temporal tiling** (`temporal_tiled_encode`): Segments the video along the time axis, optionally nesting spatial tiling within each temporal chunk for long video sequences.

Each tile is encoded independently, producing latent representations that are later stitched together.

### Seamless Blending

To eliminate visible seams between tiles, the implementation uses linear interpolation functions defined in lines 56-78 of [`dc_ae.py`](https://github.com/NVlabs/Sana/blob/main/dc_ae.py):

- `blend_v` – Blends vertical overlaps
- `blend_h` – Blends horizontal overlaps  
- `blend_t` – Blends temporal overlaps

These functions apply the `tile_overlap_factor` (default 25%) to create smooth transitions between adjacent tiles, ensuring the final latent map appears seamless despite being processed in chunks.

### The Decoding Pipeline

The decoder mirrors the encoder's approach through `spatial_tiled_decode` and `temporal_tiled_decode` (lines 160-210). Each latent tile is decoded independently, and the resulting pixel-space tiles are blended using the same overlap functions before reconstruction. This symmetry ensures that memory usage remains constrained during both compression and decompression phases.

## Enabling Tiling in Your Code

The model builder utility automatically invokes tiling when the configuration supports it. In [`diffusion/model/builder.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/model/builder.py) (line 352), the system calls:

```python
ae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)

```

For manual instantiation, construct the VAE with a tiling-enabled configuration:

```python
from diffusion.model.dc_ae.efficientvit.models.efficientvit.dc_ae import dc_vae_f32, DCAE

# Build configuration with tiling enabled

cfg = dc_vae_f32(
    name="dc-vae-f32t4c128",
    pretrained_path="hf://Efficient-Large-Model/SANA-Video_2B_480p/checkpoints/SANA_Video_2B_480p.pth",
)  # Helper sets video.use_spatial_tiling=True and video.use_temporal_tiling=True

# Instantiate model

vae = DCAE(cfg)

# Optional: Customize tile dimensions for your GPU

vae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)

```

## 4K Resolution Inference Workflow

The following implementation demonstrates processing a 3840×2160 video while maintaining memory usage appropriate for 8-GB GPUs:

```python

# 4k_tiling_demo.py

import torch
import imageio
from diffusion.model.dc_ae.efficientvit.models.efficientvit.dc_ae import dc_vae_f32, DCAE

# ----------------------------------------------------------------------

# 1. Load 4K video (3840×2160) as (T, C, H, W) tensor

# ----------------------------------------------------------------------

video_path = "sample_4k.mp4"
reader = imageio.get_reader(video_path)
frames = [
    torch.from_numpy(frame).permute(2, 0, 1).float() / 255.0 
    for frame in reader
]
video = torch.stack(frames, dim=0)  # Shape: (T, 3, 2160, 3840)

# ----------------------------------------------------------------------

# 2. Initialize VAE with tiling enabled

# ----------------------------------------------------------------------

cfg = dc_vae_f32(
    name="dc-vae-f32t4c128",
    pretrained_path="hf://Efficient-Large-Model/SANA-Video_2B_480p/checkpoints/SANA_Video_2B_480p.pth",
)
vae = DCAE(cfg)

# Force specific tile size (optional, default is 1024)

vae.enable_tiling(tile_sample_min_height=1024, tile_sample_min_width=1024)

# ----------------------------------------------------------------------

# 3. Encode with automatic tiling (memory-efficient)

# ----------------------------------------------------------------------

video = video.unsqueeze(0)  # Add batch dimension: (1, T, 3, 2160, 3840)

latent = vae.encode(video)  # Automatic spatial/temporal tiling applied

# ----------------------------------------------------------------------

# 4. Decode back to pixel space (also tiled)

# ----------------------------------------------------------------------

recon_video = vae.decode(latent)

print(f"Latent shape: {latent.shape}")           # e.g., (1, 4, 40, 68, 120)

print(f"Reconstruction shape: {recon_video.shape}")  # (1, T, 3, 2160, 3840)

```

**Key implementation details:**

- The full 4K frame never resides in GPU memory simultaneously; only individual 1024×1024 tiles are processed at any moment
- With the default compression ratio of 32, each spatial tile compresses to a 32×32 latent patch, further reducing memory overhead
- The pipeline in [`app/sana_video_refiner_pipeline_diffusers.py`](https://github.com/NVlabs/Sana/blob/main/app/sana_video_refiner_pipeline_diffusers.py) (line 99) provides a production example of manually enabling tiling before inference

## Memory Optimization Mechanism

The DC-AE tiling system achieves **constant memory scaling** regardless of input resolution. By processing tiles independently, peak GPU memory consumption is bounded by:

```

Memory_max = Memory_encoder(Tile_size) + Memory_overlap_buffer

```

For a 1024×1024 tile with 3 channels and float32 precision, this requires approximately 12 MB for the input tile, compared to 96 MB for a full 4K frame—a reduction that enables 4K video processing on hardware with limited VRAM. The latent representation receives proportionate savings, with each tile compressing to 32×32×4 channels at approximately 16 KB per tile.

## Summary

- **DC-AE tiling** in NVlabs/Sana automatically splits high-resolution content into manageable tiles without manual preprocessing
- **Configuration flags** `video.use_spatial_tiling` and `video.use_temporal_tiling` in `DCVAEConfig` control the behavior
- **Memory usage** scales with tile size (default 1024×1024) rather than frame resolution, enabling 4K inference on 8-GB GPUs
- **Seamless reconstruction** is achieved through linear blending functions (`blend_v`, `blend_h`, `blend_t`) that handle 25% overlap regions
- **Implementation** requires only calling `vae.enable_tiling()` or using predefined configurations like `dc_vae_f32()`

## Frequently Asked Questions

### What is the default tile size for DC-AE tiling in Sana?

The default configuration uses **1024×1024 pixels** for spatial tiles and 1024 frames for temporal segments. These values are set in the `dc_vae_f32` configuration helper and can be overridden by passing `tile_sample_min_height` and `tile_sample_min_width` to the `enable_tiling()` method.

### How does DC-AE prevent visible seams between tiles?

The system applies linear interpolation across overlapping regions using dedicated blending functions (`blend_v`, `blend_h`, and `blend_t`) defined in [`diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/model/dc_ae/efficientvit/models/efficientvit/dc_ae.py). With a default `tile_overlap_factor` of 0.25 (25%), adjacent tiles blend smoothly to produce a seamless final output.

### Can DC-AE tiling be used for still images as well as video?

Yes. While the configuration includes `video.use_spatial_tiling` and `video.use_temporal_tiling` flags, spatial tiling functions independently for single images (batched as `(B, C, H, W)`). Temporal tiling only activates when processing video tensors with a time dimension `(B, T, C, H, W)`.

### What is the minimum GPU memory required for 4K inference with DC-AE?

The system can process 4K (3840×2160) content on GPUs with **8 GB of VRAM** by using the default 1024×1024 tile size. Memory usage remains constant regardless of frame dimensions because the encoder and decoder process only one tile at a time, with peak consumption determined solely by the tile dimensions and model width.