Understanding the YuE Truncated Flag: Token Limits in SymbolicPlan and SemanticResult

The truncated flag on SymbolicPlan and SemanticResult indicates that a generation stage stopped because the model consumed its maximum token budget (max_tokens) before completing the full musical sequence.

The truncated boolean flag in the YuE music generation pipeline signals when token budgets constrain output length. In the multimodal-art-projection/YuE repository, both the ABC planning and semantic generation stages expose this flag to mark incomplete sequences. Developers use this indicator to diagnose generation limits and adjust GenerationConfig parameters for complete musical outputs.

What the Truncated Flag Indicates

The truncated property appears on two critical result objects in src/yue2/pipeline.py. When True, it confirms that the backend stopped token generation due to budget exhaustion rather than natural completion.

SymbolicPlan.truncated

The SymbolicPlan.truncated flag becomes True when the ABC planning step (plan()) generates tokens until reaching the max_tokens threshold defined in the ABC Sampling configuration. According to src/yue2/protocol.py, the default limit is 4096 tokens for ABC generation.

When this flag is set, the ABC transcription is incomplete. Downstream stages including semantic generation and audio synthesis will still execute, but the final song may lack musical sections that would have appeared with a higher token allowance.

SemanticResult.truncated

The SemanticResult.truncated flag activates when the semantic generation step (generate_semantic()) exhausts the max_tokens budget in the semantic Sampling config. The default semantic limit is 9000 tokens as defined in the protocol.

A True value here indicates the codec token stream is truncated. Consequently, the subsequent NAR-to-audio stage may synthesize a shorter audio clip or cut off abruptly.

How the Pipeline Sets the Truncated Flag

The flag originates in the low-level generation helper and propagates through the pipeline stages.

Token Generation Backend

In src/yue2/pipeline.py, the _generate() method calls either generate_tokens() or generate_vllm(). These backends return a tuple containing the generation status:

ids, timing, truncated = self._generate(...)

The third element (truncated) is a boolean that is True when the backend stopped specifically because the token budget was consumed. Lines 51-53 of pipeline.py show this value being captured and passed into the result constructors.

Result Object Construction

Both SymbolicPlan and SemanticResult accept the truncated parameter in their constructors. The pipeline instantiates these objects immediately after generation, preserving the truncation state for downstream logic.

Progress Reporting

The Progress.complete() method in src/yue2/progress.py (lines 37-49) uses the flag to emit user-visible status. When truncated is True, the CLI displays "Finished (generation limit reached)" instead of "Completed", providing immediate feedback that the output may be incomplete.

Practical Implications of Truncation

Understanding truncation helps developers build robust applications around the YuE pipeline.

Conditional Logic and Exit Codes

Automation scripts check the flag to determine pipeline success. In skills/yue2-music/scripts/run_yue2.py (line 147), the code examines result["truncated"] to set a non-zero exit code when truncation occurs, enabling CI/CD pipelines to detect incomplete generations.

Artifact Persistence

When saving results via SongResult.save_artifacts(), the truncated map is persisted alongside audio files. This ensures reproducibility by recording whether the generation hit token limits, allowing exact recreation of the generation state.

Code Examples

Checking Truncation Status After Generation

Inspect both flags after running the pipeline to verify output completeness:

from yue2.pipeline import YuE2Pipeline

pipe = YuE2Pipeline.from_pretrained()
result = pipe(style="Jazz", lyrics="Swing rhythm", tags="jazz")

print("ABC truncation:", result.truncated["abc"])
print("Semantic truncation:", result.truncated["semantic"])

Example output showing semantic truncation:


ABC truncation: False
Semantic truncation: True

Increasing Token Limits to Prevent Truncation

Raise the max_tokens value in the Sampling configuration to allow longer generations:

from yue2.pipeline import YuE2Pipeline, GenerationConfig, Sampling

# Increase ABC token budget from default 4096 to 8000

abc_cfg = Sampling(max_tokens=8000)
gen_cfg = GenerationConfig(abc=abc_cfg)

pipe = YuE2Pipeline.from_pretrained(generation_config=gen_cfg)
result = pipe(style="Classical", lyrics="Moonlight", tags="classical")

print(result.truncated)  # Likely {"abc": False, "semantic": False}

Using Truncation Flags in CLI Scripts

Integrate truncation detection into automated workflows:

import json
from pathlib import Path
from yue2.pipeline import YuE2Pipeline

pipe = YuE2Pipeline.from_pretrained()
result = pipe(style="Pop", lyrics="Love song", tags="pop")
output_dir = Path("out")
result.save_artifacts(output_dir)

summary = {
    "truncated": result.truncated,
    "audio_seconds": len(result.audio) / result.sample_rate,
}
print(json.dumps(summary, indent=2))

Key Source Files

The truncation mechanism spans several files in the YuE codebase:

Summary

  • The truncated flag on SymbolicPlan and SemanticResult indicates the model reached its max_tokens limit before completing generation.
  • ABC planning defaults to 4096 tokens while semantic generation defaults to 9000 tokens, defined in src/yue2/protocol.py.
  • The flag propagates from the _generate() method through result objects to progress reporters and CLI outputs.
  • Automation scripts use this flag to set error codes and trigger retries with increased token budgets.
  • Truncation state persists in saved artifacts via SongResult.save_artifacts() for reproducibility.

Frequently Asked Questions

What causes the truncated flag to become True in YuE?

The truncated flag becomes True when either the ABC planning stage or semantic generation stage exhausts its allocated token budget. This occurs when the number of generated tokens reaches the max_tokens parameter in the respective Sampling configuration before the model naturally completes the sequence.

How can I prevent truncation in YuE generations?

Increase the max_tokens value in your GenerationConfig. For ABC planning, raise it above the default 4096 tokens. For semantic generation, raise it above the default 9000 tokens. Note that higher limits increase memory usage and generation time but produce more complete musical outputs.

Does truncation affect the final audio quality?

Truncation affects completeness rather than per-token quality. If SymbolicPlan is truncated, the song structure may be missing sections. If SemanticResult is truncated, the audio may end abruptly or be shorter than intended. The synthesized audio itself remains high quality for the tokens that were generated.

Where is the truncated flag stored after generation?

The truncated flag is stored in the SongResult object returned by the pipeline. When you call result.save_artifacts(), the truncation state persists alongside the audio files and metadata, ensuring you can later determine whether a generation completed naturally or hit token limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →