# How YuE's `decode_tiled` Splits Acoustic Latents: Core Tiling vs Full Decoding

> Understand YuE's decode_tiled: it splits acoustic latents into core tiles with halo regions or uses full decoding. Learn when each method is best for your data.

- Repository: [multimodal-art-projection/YuE](https://github.com/multimodal-art-projection/YuE)
- Tags: internals
- Published: 2026-09-14

---

**The `decode_tiled` method in [`src/yue2/modeling_vae.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/modeling_vae.py) splits long acoustic latents into independent core tiles padded with halo regions to preserve decoder context, while full decoding processes entire tensors in one pass for shorter sequences.**

The YuE music generation framework provides two distinct paths for converting acoustic VAE latents into waveforms: a memory-efficient tiled approach and a full-frame approach. Understanding how `decode_tiled` segments latent tensors—and when to prefer the full decoding path—is essential for optimizing GPU memory usage and inference latency in the `multimodal-art-projection/YuE` codebase.

## How `decode_tiled` Splits Acoustic Latents

The `decode_tiled` implementation (lines 39‑79 of [`src/yue2/modeling_vae.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/modeling_vae.py)) processes latent tensors by dividing them into overlapping segments that maintain exact acoustic context.

### Core and Halo Division

The algorithm accepts `core_frames` (default `decode_core_frames = 1024`) defining the output segment size and `halo_frames` (default `decode_halo_frames = 16`) providing receptive-field padding. For each tile, the code calculates expanded boundaries to capture convolutional dependencies:

```python
left = max(0, start - halo_frames)      # line 67

right = min(frames, end + halo_frames)  # line 68

tile = self.decode(latent[..., left:right])  # line 69

```

By decoding the expanded slice `latent[..., left:right]`, the VAE ensures that every convolutional operation affecting the core segment has access to its required context window.

### Exact-Core Cropping

After decoding the expanded tile, the implementation crops the output to the exact core length using the model's downsampling ratio (`self.config.downsampling_ratio`):

```python
out_start, out_end = start * ratio, min(end * ratio, total)      # line 70

crop_start = (start - left) * ratio                               # line 71

crop = tile[..., crop_start:crop_start + out_end - out_start]    # line 72

```

This cropping guarantees that the final waveform is **numerically identical** (within floating-point rounding) to the output of full decoding, as the halos only provide temporary context rather than contributing to the final output boundaries.

### Progress Reporting and Memory Management

The method includes an optional `on_progress(completed, total)` callback invoked after each tile (lines 77‑78), enabling real-time progress tracking for long sequences. By default, `output_device="cpu"` assembles the final waveform on the host, keeping GPU memory usage constant regardless of audio length.

## When to Use Full Decoding vs `decode_tiled`

The choice between decoding strategies depends on sequence length, memory constraints, and latency requirements:

| Situation | Preferred Method | Reason |
|-----------|------------------|---------|
| **Short audio** (latents < `decode_core_frames`) | **Full decode** (`decode`) | Eliminates tiling overhead when the entire sequence fits in a single core frame. |
| **Limited GPU RAM** | **Tiled decode** (`decode_tiled`) | Processes tiles sequentially and stores intermediates on CPU, preventing out-of-memory errors. |
| **Streaming playback** | **Tiled decode** | Enables chunked output as tiles complete, reducing time-to-first-audio. |
| **Maximum fidelity** | **Either** | Both paths produce identical waveforms; select based on resource constraints. |

Full decoding (lines 21‑26 of [`modeling_vae.py`](https://github.com/multimodal-art-projection/YuE/blob/main/modeling_vae.py)) bypasses tiling entirely, forwarding the complete latent tensor to the decoder in one call without halo calculations or cropping logic.

## Pipeline Integration and Automatic Selection

The high-level pipeline in [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) automates path selection based on the `full` parameter:

```python
if full:
    audio = model.decode(z.to(self.device)).cpu()                  # line 45

else:
    audio = model.decode_tiled(z, core_frames=self.vae_core_frames,
                               halo_frames=16, output_device="cpu",
                               on_progress=report)                # lines 48-49

```

Users can force full decoding by passing `full=True` to `Pipeline.decode()`, while the default behavior uses tiled decoding for long sequences.

## Practical Code Examples

### Basic Tiled Decoding

Explicitly invoke `decode_tiled` with default parameters for long latent sequences:

```python
from yue2.modeling_vae import YuE2VAE
import torch

vae = YuE2VAE.from_pretrained("example/decoder", decoder_only=True, device="cpu")
latent = torch.randn(1, 64, 5000)  # [B, latent_dim, T]

# 1024-frame cores with 16-frame halos

audio = vae.decode_tiled(latent)  # → Tensor[B, 2, ~9.6M samples]

```

### Full Decoding (No Tiling)

Process short latents in a single pass without tiling overhead:

```python
audio_full = vae.decode(latent)  # Identical output, no halo padding

```

### Automatic Selection via Pipeline

Allow the pipeline to choose the optimal decoding strategy:

```python
from yue2.pipeline import Pipeline

pipe = Pipeline(model="example/song-model", vae_dir="example/decoder")
latents = pipe.synthesize(...)  # Generate latent representation

# Short audio: forces full decode

short_wave = pipe.decode(latents, full=True)

# Long audio: automatically uses tiled decode (default)

long_wave = pipe.decode(latents)

```

### Progress Callbacks for Streaming

Monitor decoding progress for user feedback or streaming applications:

```python
def report_progress(completed: int, total: int):
    print(f"Decoded tile {completed}/{total}")

audio = vae.decode_tiled(latent, on_progress=report_progress)

```

## Summary

- **`decode_tiled`** in [`src/yue2/modeling_vae.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/modeling_vae.py) splits acoustic latents into overlapping core tiles with configurable halo regions to maintain decoder context.
- **Exact-core cropping** ensures tiled outputs match full-decoding fidelity by removing halo contributions after processing.
- **Full decoding** is preferable for short sequences that fit within a single core tile, avoiding unnecessary tiling overhead.
- **Tiled decoding** is required for long sequences or memory-constrained environments, processing audio incrementally while keeping GPU usage constant.
- The **pipeline** automatically selects the appropriate path unless overridden with `full=True`.

## Frequently Asked Questions

### What is the purpose of halo frames in `decode_tiled`?

Halo frames provide the acoustic context required by the VAE decoder's convolutional layers. According to the implementation in [`src/yue2/modeling_vae.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/modeling_vae.py), the 16-frame default ensures that receptive fields extending beyond tile boundaries can access necessary latent information without introducing boundary artifacts into the core output.

### Does tiled decoding affect audio quality compared to full decoding?

No. The `decode_tiled` implementation performs exact-core cropping using the downsampling ratio to ensure the final waveform is numerically identical to full decoding output. Both paths produce the same acoustic result; tiled decoding merely trades compute efficiency for memory efficiency.

### How does the YuE pipeline decide between tiled and full decoding?

The pipeline in [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) relies on the explicit `full` boolean parameter rather than automatic length detection. When `full=True`, it calls `model.decode()` directly; otherwise, it invokes `decode_tiled()` with configured `core_frames` and `halo_frames` parameters.

### Can I adjust the tile size for memory-constrained GPUs?

Yes. The `core_frames` parameter controls tile size and can be reduced below the default 1024 to lower peak GPU memory usage. However, smaller tiles increase CPU-GPU transfer overhead and callback frequency, so values below 512 are generally not recommended unless absolutely necessary for memory constraints.