How `save_artifacts` Lays Out Output Files for Full Run Reproducibility in YuE

The save_artifacts method in src/yue2/pipeline.py serializes the semantic plan, audio output, and intermediate representations to a single directory, capturing every state needed to recreate an identical generation.

The multimodal-art-projection/YuE repository implements a complete audio generation pipeline where reproducibility is critical for research and production workflows. When you invoke save_artifacts, it packages the exact conditions that produced your audio into a portable folder structure. This ensures that anyone with the same model weights can reconstruct the identical output without guessing parameters or reprocessing prompts.

The Four-File Reproducibility Bundle

Inside the target directory, save_artifacts creates four distinct files that work together to freeze the generation state. Each file serves a specific purpose in the reconstruction chain, from high-level plan metadata to low-level decoder latents.

  • pipeline.json: Generated by self.semantic.plan.save(), this JSON file captures the semantic plan—including generation steps, token IDs, random seeds, and mutable parameters that directed the model's behavior.
  • audio.flac: The final rendered audio file produced by self.save(), providing the immediate output and a reference for verification.
  • semantic.npy: A NumPy array (dtype int32) containing the raw semantic token IDs (self.semantic.tokens), representing the intermediate symbolic representation of the audio.
  • latent.npy: A NumPy array (dtype float32) storing the latent vectors (self.latents) fed into the decoder, preserving the exact internal state needed for deterministic reconstruction.

Implementation Details in src/yue2/pipeline.py

The save_artifacts method is implemented between lines 103 and 110 of src/yue2/pipeline.py. The method accepts a directory parameter and writes the four artifacts in a strict sequence to ensure data consistency.

def save_artifacts(self, directory):
    # Store the plan that drove the generation

    self.semantic.plan.save(directory)
    # Write the final audio file

    self.save(directory / "audio.flac")
    # Persist the semantic token sequence

    np.save(directory / "semantic.npy",
            np.asarray(self.semantic.tokens, dtype=np.int32))
    # Persist the latent vectors used by the decoder

    np.save(directory / "latent.npy", self.latents.astype(np.float32))

Why This Layout Guarantees Reproducibility

The combination of these four files creates a deterministic snapshot of the generation process. By separating concerns between high-level planning and low-level execution, YuE enables both human inspection and machine-perfect replication.

Plan JSON captures the exact sequence of operations and any stochastic elements like random seeds or temperature settings. Without this metadata, reconstructing the conditions of the generation would be impossible.

Semantic tokens (semantic.npy) store the discrete symbolic representation generated by the language model component. These tokens form the bridge between the text prompt and the final audio, allowing you to skip the expensive autoregressive generation step during reproduction.

Latent vectors (latent.npy) preserve the continuous representations passed to the audio decoder. Since these are the actual inputs to the vocoder, reloading them guarantees that the decoder produces bit-identical audio waveforms.

Audio file serves as the ground truth artifact, allowing immediate verification that a reproduction attempt matches the original output exactly.

Working with save_artifacts Programmatically

Saving a Generation Run

To create a reproducible artifact bundle, instantiate the Pipeline class, run your generation, and call save_artifacts with a target path.

from yue2.pipeline import Pipeline
from pathlib import Path

# Initialize the pipeline from pretrained weights

pipe = Pipeline.from_pretrained("yue2-base")

# Execute generation with your specific parameters

result = pipe(prompt="A calm evening piano piece", duration_seconds=30)

# Persist the complete run state

output_dir = Path("my_reproducible_run")
result.save_artifacts(output_dir)

Reloading and Verifying a Saved Run

You can restore the exact generation state by manually loading the plan and NumPy arrays back into a fresh pipeline instance.

from yue2.pipeline import Pipeline
import numpy as np
from pathlib import Path

# Initialize a fresh pipeline with identical model weights

pipe = Pipeline.from_pretrained("yue2-base")

# Load the saved artifacts

saved_path = Path("my_reproducible_run")
pipe.semantic.plan.load(saved_path)  # Reads pipeline.json

pipe.semantic.tokens = np.load(saved_path / "semantic.npy")
pipe.latents = np.load(saved_path / "latent.npy")

# Regenerate the audio (should match original byte-for-byte)

pipe.save(saved_path / "recreated.flac")

Integration with the YuE Ecosystem

The save_artifacts method relies on several supporting components throughout the codebase.

src/yue2/protocol.py defines the Plan class that handles serialization of pipeline.json. This class manages the semantic plan's schema, ensuring backward compatibility when reloading older generation runs.

src/yue2/cli.py serves as the command-line entry point that orchestrates the save process. When running YuE from the terminal, the CLI invokes plan.save followed by result.save_artifacts to persist both the configuration and results.

src/yue2/storage.py provides underlying utilities for checkpoint management. While save_artifacts handles individual run states, the storage module ensures that target directories are properly managed and that no checkpoint collisions occur during save_pretrained operations.

tests/test_progress_integration.py contains unit tests that verify save_artifacts invocation. The test suite uses mock implementations (save_artifacts=lambda directory: {'identity': 'saved'}) to ensure the pipeline correctly triggers artifact serialization without requiring full model inference during testing.

Summary

  • save_artifacts in src/yue2/pipeline.py creates a four-file bundle that completely captures a generation run's state.
  • The output directory contains pipeline.json (semantic plan), audio.flac (final output), semantic.npy (token IDs), and latent.npy (decoder inputs).
  • This layout enables full reproducibility by preserving both high-level generation parameters and low-level latent representations.
  • You can reload these artifacts into a fresh Pipeline instance to recreate bit-identical audio without rerunning the expensive generation phase.
  • The method integrates with src/yue2/protocol.py for plan serialization and is tested in tests/test_progress_integration.py.

Frequently Asked Questions

What file format does YuE use for the semantic plan?

YuE serializes the semantic plan as pipeline.json inside the artifact directory. This JSON file is generated by the save method of the Plan class defined in src/yue2/protocol.py, capturing generation steps, token IDs, and random seeds in a human-readable schema.

Can I reproduce audio without re-running the full model?

Yes. By loading semantic.npy and latent.npy into a fresh pipeline instance, you can bypass the autoregressive language model generation and proceed directly to the decoding phase. This works because the latent vectors represent the exact inputs that the decoder requires to synthesize audio.

Where is the save_artifacts method defined?

The method is defined in src/yue2/pipeline.py between lines 103 and 110. It belongs to the main Pipeline class and orchestrates calls to self.semantic.plan.save, self.save, and NumPy's save function to persist the four reproducibility files.

Does save_artifacts overwrite existing files?

The method writes directly to the supplied directory using standard file operations. According to the implementation in src/yue2/pipeline.py, it does not explicitly check for existing files before calling np.save or self.save, meaning it will overwrite any pre-existing audio.flac, semantic.npy, or latent.npy files in the target location. Ensure you specify a unique or empty directory to prevent data loss.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →