# What the legacy_off BF16-vs-FP32 Logits Branch Preserves from YuE1 Decoding

> Discover how the legacy_off BF16-vs-FP32 logits branch maintains bitwise compatibility by preserving BF16-quantized logits from YuE1 decoding, avoiding FP32 conversion for precise token selection in YuE2.

- Repository: [multimodal-art-projection/YuE](https://github.com/multimodal-art-projection/YuE)
- Tags: internals
- Published: 2026-09-14

---

**The legacy_off BF16-vs-FP32 logits branch preserves the original BF16-quantized logits from the YuE1 inference pipeline, storing the low-precision 16-bit representations instead of converting them to full-precision FP32 to ensure bitwise-compatible token selection in YuE2.**

The **multimodal-art-projection/YuE** repository includes this specialized branch to maintain exact numerical fidelity with the original YuE1 decoding behavior. By retaining the native **BF16 (BFloat16)** logits rather than automatically upcasting to FP32, the branch guarantees that the newer YuE2 pipeline reproduces identical sampling decisions and generation results to the legacy implementation.

## Preserving BF16 Logits from YuE1 Decoding

The **legacy_off BF16-vs-FP32 logits** branch maintains the original low-precision representations that the YuE1 decoder produced during token generation. Instead of performing an automatic type conversion to FP32—which introduces subtle numerical differences—the branch caches and reuses the exact BF16 values.

This preservation ensures **numerical consistency** across model versions. Since sampling algorithms like top-k and nucleus sampling operate directly on logit values, even minor precision changes alter the probability distribution and ultimately change which tokens get selected. The branch eliminates this drift by serving the original BF16 tensors to the sampling logic.

## Why Numerical Fidelity Matters

### Reproducible Sampling Decisions

By keeping logits in their original BF16 format, the branch ensures that stochastic sampling decisions remain identical to those made by YuE1. When the YuE2 pipeline activates the preserved logits path via the **`preserve_logits=True`** flag, it bypasses the default FP32 conversion that standard inference uses.

### Benchmark Decoder Compatibility

The official evaluation decoder, **YuE2-Vae-legacy**, expects the exact BF16 logits produced by the original implementation. Converting these values to FP32 introduces numerical drift that can affect reported scores in quantitative benchmarks. Preserving the BF16 representation ensures that FAD, CLAP, and other automated metrics remain comparable across YuE1 and YuE2 evaluations.

### Research Utility

The branch provides a controlled environment for studying **low-precision versus full-precision logits** without modifying underlying model weights. Researchers can toggle between BF16 and FP32 modes to quantify how 16-bit quantization affects musical coherence, style adherence, and prompt alignment, all while maintaining the ability to fall back to YuE1-equivalent behavior.

## Implementation in the YuE2 Pipeline

The preservation mechanism is implemented across several core modules in the codebase.

### Key Source Files

- **[`src/yue2/quantization.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/quantization.py)**: Implements the BF16 quantizer and the logic that saves original logits for later reuse. This module handles the low-level bit representation storage and retrieval.
- **[`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py)**: Contains the core inference pipeline where the **`preserve_logits`** flag routes decoding through the stored BF16 logits instead of converting to FP32.
- **[`src/yue2/modeling_yue2.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/modeling_yue2.py)**: Defines the model architecture where logits are initially produced in BF16 before any precision conversion occurs.
- **[`docs/generation.md`](https://github.com/multimodal-art-projection/YuE/blob/main/docs/generation.md)**: Documents the "legacyoff" mode and explains the methodology for enabling preserved logits during generation.

### Activating Preserved Logits

To access the preserved BF16 logits, instantiate the pipeline using the legacyoff branch model and set the preservation flag:

```python
from yue2.pipeline import YuE2Pipeline

# Load the legacy-off branch model

pipeline = YuE2Pipeline.from_pretrained(
    "m-a-p/YuE2-3B-legacyoff-bfi6-fp32", 
    device="cuda"
)

# Decode using the preserved BF16 logits

song = pipeline(
    lyrics="When sunrise paints the sky",
    style="Acoustic folk",
    cot="full",
    preserve_logits=True,  # Forces use of stored BF16 logits

)
song.save_artifacts("outputs/legacy_preserved")

```

### Comparing BF16 vs. FP32 Logits

Researchers can extract both precision representations to study their divergence:

```python
from yue2.quantization import BF16Quantizer, FP32Quantizer

# Compare logits in both precision modes

logits_bf16 = pipeline.decode_logits(mode="bf16")
logits_fp32 = pipeline.decode_logits(mode="fp32")

# Quantify the numerical difference

max_diff = (logits_bf16 - logits_fp32).abs().max()
print(f"Max absolute difference: {max_diff}")

```

This comparison reveals how precision loss in the BF16 path affects the final probability distribution before sampling.

## Summary

- The **legacy_off BF16-vs-FP32 logits** branch stores original BF16-quantized logits from YuE1 instead of converting them to FP32.
- Preserving these values ensures **bitwise-compatible token selection** and reproducible sampling decisions across YuE1 and YuE2 pipelines.
- The **`preserve_logits=True`** parameter in [`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py) activates the preserved BF16 path during inference.
- Benchmark evaluations using **YuE2-Vae-legacy** require these exact logits to maintain valid comparison metrics.
- The implementation in [`src/yue2/quantization.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/quantization.py) enables research into low-precision versus full-precision generation quality.

## Frequently Asked Questions

### What is the difference between the legacy_off branch and standard YuE2 inference?

Standard YuE2 inference automatically converts logits to FP32 for numerical stability, while the legacy_off branch retains the original BF16 representations produced by YuE1. This ensures backward-compatible results but requires explicit activation via the `preserve_logits` parameter in the pipeline configuration.

### Why does preserving BF16 logits matter for benchmark results?

Sampling algorithms like top-k and nucleus sampling are sensitive to small numerical differences in the logit distribution. Converting BF16 to FP32 changes these values slightly, which can alter token selection chains and impact quantitative metrics such as FAD or CLAP scores reported in official evaluations.

### Can I switch between BF16 and FP32 logits during the same session?

Yes. The pipeline supports dynamic mode switching through the `decode_logits()` method, allowing researchers to extract both representations for comparison without reloading the model weights. Use `mode="bf16"` to access the preserved low-precision logits or `mode="fp32"` to see the converted full-precision values.

### Where is the BF16 quantization logic implemented?

The quantizer and storage logic resides in **[`src/yue2/quantization.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/quantization.py)**, which handles the saving and loading of low-precision logits. The decision to use these cached values versus fresh FP32 conversions occurs in **[`src/yue2/pipeline.py`](https://github.com/multimodal-art-projection/YuE/blob/main/src/yue2/pipeline.py)** when the `preserve_logits` flag is set to `True`.