How the `negative_prefix` CFG Branch Differs from Positive Prefix in YuE2 and Why Symbolic CFG Retains Exact ABC IDs
In the YuE2 protocol, the negative_prefix branch omits ABC tokens for audio-only generation but must replicate the exact same ABC IDs as the positive branch during symbolic CFG to prevent the classifier-free guidance subtraction from corrupting the symbolic score.
The multimodal-art-projection/YuE repository implements a specialized generation pipeline where classifier-free guidance (CFG) operates through two distinct prefix constructions. Understanding how the negative_prefix() function builds its token stream compared to the standard token_prefixes() function is essential for maintaining ABC notation integrity during music generation tasks.
Structural Differences Between CFG Branches
When CFG is enabled, YuE2 runs the model twice and combines the logits using the formula logits = pos_logits + cfg_scale * (pos_logits – neg_logits). The two branches receive different token prefixes depending on the chain-of-thought (CoT) mode:
Positive (prompt) prefix
- Constructed by
token_prefixes(request, tokenizer, abc_ids)insrc/yue2/protocol.py - Sequence:
EODtoken + encoded instruction text (INSTRUCTIONS[request.cot]) + optional ABC token IDs +ABC_START,ABC_END,MUSIC_START
Negative (guidance) prefix
- Constructed by
negative_prefix(request, tokenizer, abc_ids)insrc/yue2/protocol.py - Sequence:
EODtoken + instruction text only +MUSIC_START(whencot=="off") or the identical ABC token IDs wrapped withABC_START/ABC_END(when symbolic CFG is active)
Positive Prefix Implementation
In src/yue2/protocol.py, the positive branch encodes the full user request including any symbolic score data:
def token_prefixes(request, tokenizer, abc_ids=None):
base = [EOD] + tokenizer.encode(request.text())
# … handle "off", "melody", "full" …
return base + [ABC_START] + abc_ids + [ABC_END, MUSIC_START]
This function returns the complete context needed for generation, including the ABC notation tokens when present.
Negative Prefix Implementation
The negative branch strips away the user's content but retains structural markers. According to the source code in src/yue2/protocol.py:
def negative_prefix(request, tokenizer, abc_ids=None):
base = [EOD] + tokenizer.encode(INSTRUCTIONS[request.cot])
if request.cot == "off":
return base + [MUSIC_START]
if abc_ids is None:
raise ValueError("Symbolic CFG must retain the exact positive‑branch ABC IDs")
return base + [ABC_START] + abc_ids + [ABC_END, MUSIC_START]
When cot is set to "off" (audio-only), the function returns immediately after adding MUSIC_START. For symbolic generation modes ("melody" or "full"), the function validates that abc_ids is provided and inserts the exact same token sequence used in the positive branch.
Why Symbolic CFG Requires Exact ABC ID Retention
Standard CFG works by subtracting the negative prediction from the positive prediction to amplify the signal of the desired content. However, this mathematical operation assumes that both branches share the same structural context except for the guided attribute.
If the negative branch omitted or altered the ABC token IDs during symbolic generation, the subtraction pos_logits – neg_logits would treat the symbolic score itself as a difference to be amplified or suppressed. This would corrupt the exact note identifiers, rhythms, and bar lines that define the ABC notation format.
By requiring the negative branch to use the identical abc_ids array via the negative_prefix() validation logic, YuE2 ensures that the CFG transformation only influences the semantic and acoustic rendering of the music while leaving the symbolic representation mathematically neutral. The ValueError raised in src/yue2/sampling.py when abc_ids is missing enforces this constraint, guaranteeing that the generated ABC file remains identical to the input score (apart from any explicit user edits).
Practical CFG Implementation Example
The following example demonstrates how the two prefixes are constructed and used within the YuE2 pipeline when processing a symbolic request:
from yue2.protocol import token_prefixes, negative_prefix
from yue2.pipeline import YuE2Pipeline
# Build a request with symbolic score (cot="full")
request = {
"style": "pop, bright synth",
"lyrics": "Hey, we dance all night...",
"cot": "full",
"abc": "X:1\nT:Demo\nM:4/4\nK:C\nC D E F|G A B c|",
"cfg_scale": 1.2,
}
# Initialize pipeline and construct prefixes
with YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cpu") as pipe:
tokenizer = pipe.tokenizer
# Positive prefix with full ABC content
pos_prefix = token_prefixes(request, tokenizer)
# Negative prefix must reuse identical ABC IDs
abc_ids = tokenizer.encode(request["abc"])
neg_prefix = negative_prefix(request, tokenizer, abc_ids=abc_ids)
# Pipeline internally executes both branches and applies CFG scaling
song = pipe(**request)
When cot="off", the call to negative_prefix() returns [EOD] + instruction + [MUSIC_START] without requiring ABC tokens. For "melody" or "full" modes, omitting the abc_ids parameter triggers the validation error to prevent symbolic corruption.
Summary
- The
negative_prefixfunction insrc/yue2/protocol.pyconstructs the CFG guidance stream by including only instruction tokens for audio-only mode, but requires exact ABC ID replication for symbolic generation. - When
cotis"off", the negative branch skips ABC tokens entirely; whencotis"melody"or"full", it validates and reuses the identicalabc_idsfrom the positive branch. - The
ValueError("Symbolic CFG must retain the exact positive‑branch ABC IDs")enforces exact ID retention to prevent the CFG subtraction formula from corrupting the symbolic score. - The orchestration logic in
src/yue2/pipeline.pyexecutes the two-pass inference and appliescfg_scaleto the combined logits.
Frequently Asked Questions
What happens if abc_ids is None when using symbolic CFG?
The negative_prefix() function raises ValueError("Symbolic CFG must retain the exact positive‑branch ABC IDs"). This validation prevents the model from computing CFG logits with mismatched symbolic contexts, which would otherwise corrupt the ABC notation during the subtraction step.
How does YuE2's CFG implementation differ from standard diffusion models?
Standard diffusion CFG typically uses an empty or null prompt for the negative branch. YuE2 instead uses a structured negative prefix that preserves instruction tokens and, critically for symbolic modes, the exact ABC token IDs. This specialized approach treats the symbolic score as immutable context rather than content to be guided.
Why does the negative prefix skip ABC tokens when cot="off"?
When cot="off" (audio-only generation), there is no symbolic score to preserve. The negative branch therefore omits ABC tokens entirely, containing only the EOD marker, instruction text, and MUSIC_START. This allows the CFG mechanism to guide purely the audio characteristics without reference to symbolic notation.
Which source files contain the CFG validation logic?
The primary implementation resides in src/yue2/protocol.py, which defines both token_prefixes() and negative_prefix(). Additional validation occurs in src/yue2/sampling.py, which raises errors for missing negative prefixes, while src/yue2/pipeline.py orchestrates the dual-branch execution and logit combination.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →