How YuE's `chunk_ranges` Helper Computes NAR Chunk Sizes: Formula and Constraints

The chunk_ranges helper computes evenly-spaced acoustic chunk boundaries by halving the remaining context after accounting for the autoregressive prefix and special tokens, with the strict limiting constraint being whichever is smaller: the halved remainder or the model's hard-coded CONTEXT constant.

In the YuE music generation pipeline (multimodal-art-projection/YuE), the non-autoregressive (NAR) acoustic solver processes audio frames in discrete windows to respect transformer context limits. The chunk_ranges function in src/yue2/protocol.py determines how the acoustic sequence is partitioned into these digestible chunks, ensuring the model never receives more tokens than it can attend to. Understanding this helper is essential for debugging generation failures or optimizing memory usage when synthesizing long songs.

How chunk_ranges Computes NAR Chunk Sizes

When synthesizing a song, YuE first builds an autoregressive (AR) prefix, then processes the remaining acoustic frames (the codec). Because the NAR solver works on each chunk independently, chunk_ranges must generate start–end index pairs that respect the model’s context window while covering the full codec length.

The Size Calculation Formula

At line 42 of src/yue2/protocol.py, the chunk size is derived from the available context budget:

size = min((context - prefix_tokens - 3) // 2, CONTEXT)

This arithmetic encodes several critical constraints:

  • context – The maximum number of tokens the model can attend to (typically matching the global CONTEXT constant).
  • prefix_tokens – The length of the autoregressive prefix already generated.
  • -3 – Reserves space for mandatory start/end markers (such as MUSIC_END and delimiters).
  • // 2 – Divides the remaining budget by two, reserving roughly half for the visible acoustic chunk and half for the prefix plus required padding.
  • min(..., CONTEXT) – Clamps the result to the absolute hard limit of the model’s positional capacity.

Chunk Boundary Generation

Once the size is determined, line 45 generates the actual ranges using a list comprehension:

return [(a, min(a + size, frames)) for a in range(0, frames, size)]

This produces a list of (start, end) tuples. Starting from 0, the function steps through the total number of frames in increments of size, emitting inclusive start indices and exclusive end indices. The final chunk may be shorter if the total frame count is not an exact multiple of the computed size.

The Limiting Constraint on Acoustic Chunk Sizing

The effective NAR chunk length is bounded by the available context after accounting for the prefix. Specifically, the constraint is dual-layered:

  1. Dynamic Arithmetic Limit: (context - prefix_tokens - 3) // 2
    This represents the remaining budget after the specific prefix and delimiters are consumed. As the prefix grows, this value shrinks, forcing smaller chunks.

  2. Static Global Limit: CONTEXT
    Even if the arithmetic result exceeds the model's maximum capacity (for example, with a very short prefix), the min() guard ensures the chunk never exceeds the hard-coded constant.

The most restrictive factor prevails. Consequently, the effective limit is whichever is smaller: the remaining context after the prefix, or the global CONTEXT constant. If either calculation yields less than one token, the function raises a ValueError at line 44 ("Empty codec or prefix leaves no acoustic context"), preventing invalid forward passes.

Implementation Details and Validation

The helper performs strict validation before returning ranges. At line 44 in src/yue2/protocol.py, it checks:

if frames < 1 or size < 1:
    raise ValueError("Empty codec or prefix leaves no acoustic context")

This guarantees that both the input frames (the codec length) and the computed size are non-empty. The consumer of these ranges is src/yue2/nar.py, which uses the returned tuples to slice the acoustic tensor for flow-matching inference.

Practical Code Examples

The following snippets demonstrate how chunk_ranges adapts to different prefix lengths and context windows.

Example 1: Typical generation scenario

from yue2.protocol import chunk_ranges, CONTEXT

codec_frames = 2000          # number of acoustic frames in the codec

prefix_len = 50              # length of the AR prefix (tokens)

ctx = 1024                   # model's context window

ranges = chunk_ranges(codec_frames, prefix_len, ctx)
print(ranges)

# → [(0, 485), (485, 970), (970, 1455), (1455, 2000)]

Example 2: Large prefix constraining chunk size


# When the prefix consumes most of the context, chunks become smaller

prefix_len = 900
ranges = chunk_ranges(codec_frames, prefix_len, ctx)
print(ranges)

# → [(0, 62), (62, 124), …]  # size limited by the remaining context

Example 3: Debugging with a reduced context window


# Forcing a tiny context for debugging or low-memory scenarios

small_ctx = 256
ranges = chunk_ranges(codec_frames, prefix_len=20, context=small_ctx)
print(ranges)

# → [(0, 118), (118, 236), (236, 354), …]  # size limited by CONTEXT = 256

Summary

  • chunk_ranges in src/yue2/protocol.py splits acoustic frames into NAR-compatible windows for the flow-matching solver in src/yue2/nar.py.
  • The chunk size formula halves the remaining context after subtracting the AR prefix length and three special delimiter tokens.
  • The limiting constraint is the minimum of the dynamically computed halved remainder and the global CONTEXT constant.
  • The function validates that both the codec length and computed chunk size are positive, raising a ValueError if the prefix leaves no acoustic context.
  • It returns a list of (start, end) tuples representing evenly-spaced (except possibly the final) intervals covering the full codec.

Frequently Asked Questions

Where is the chunk_ranges function defined in YuE?

The function is defined in src/yue2/protocol.py, specifically between lines 41 and 46. This file also contains related constants such as CONTEXT and CODEC_SIZE that govern the model's token budget.

Why does the formula divide the remaining context by two?

The division by two (// 2) reserves roughly half of the available token budget for the visible acoustic chunk and the other half for the prefix context, padding, and delimiter tokens. This ensures the NAR solver maintains sufficient historical context while processing new acoustic frames.

What happens if the autoregressive prefix is too long?

If prefix_tokens consumes so much of the context that (context - prefix_tokens - 3) // 2 evaluates to less than one, the function raises a ValueError stating "Empty codec or prefix leaves no acoustic context". This prevents the NAR solver from attempting to process an impossible window size.

How does chunk_ranges interact with the NAR solver?

The chunk_ranges output is consumed by src/yue2/nar.py, which iterates over the returned (start, end) tuples to slice the acoustic codec into batches. Each batch is processed independently through the acoustic flow-matching model, ensuring no single forward pass exceeds the transformer's positional embedding limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →