Seedance 2.0 Audio Post Delivery Process: From Prompt to Broadcast Master

The Seedance 2.0 audio post delivery process generates synchronized video and audio in a single pass using explicit prompt-level audio descriptions, then extracts professional stems and metadata according to specifications defined in references/audio-post-delivery.md to ensure broadcast-ready output.

Seedance 2.0 is an open-source generative media framework that treats audio as a first-class element of every prompt. According to the Emily2040/seedance-2.0 repository, the workflow splits generation and delivery into two distinct layers: prompt-level audio conditioning that locks timing during synthesis, and post-level deliverable extraction that prepares assets for downstream editors and automated pipelines.

Prompt-Level Audio Conditioning

Seedance 2.0 requires explicit audio descriptions in prompts to synchronize generation. When you write a prompt, you describe the type of audio needed so the model can lock timing and mood during the single-pass generation process.

Audio Categories and Prompt Syntax

The model recognizes five distinct audio categories that must be described explicitly in prompts, as documented in references/audio-guide.md (lines 11-14):

  • Dialogue: Short quoted lines, one speaker per turn
  • Ambience: Room tone, traffic, rain, crowd bed, machine hum
  • SFX: One-or-two concrete sounds tied to a visible action
  • Music: Broad tempo/energy/instrumentation, or explicit "silence"
  • Sync cue: Visible event that marks a beat (e.g., "door slam at endpoint")

Example prompt structure:
Dialogue: "I'm late!" Sound: low room tone + distant rain. SFX: door slam at 3 s. Music: none until after the line.

The model generates video and audio together in a single pass, so the timing specified in the prompt becomes the "clock" for the entire clip.

The Audio Planning Template

Every prompt should include a one-line Audio Planning Template that guides the post-process. According to references/audio-post-delivery.md (lines 30-33), the template follows this syntax:

Dialogue: [speaker/line/language]. Ambience: [bed]. SFX: [visible sync cue].
Music: [tempo/mood or none]. Reference: [Audio1 role].
Post: [stems/M&E/loudness/sync notes].

When this template is parsed by the Prompt Engine (implemented in skills/seedance-audio/SKILL.md), it injects audio conditioning into the model's generation parameters.

Post-Level Audio Deliverables

After generation, you must hand off a set of stems and metadata so downstream editors can finish the project for broadcast or platform delivery.

Required Stems and Metadata

The specific deliverables are listed in references/audio-post-delivery.md (lines 19-27):

Deliverable Purpose
Full mix Final stereo/5.1/Atmos (or platform-required) mix
Dialogue stem Isolated spoken lines for dubbing, edits, cleanup
Music stem Music-only layer
Effects stem All SFX and designed sound
M&E Music + Effects without original dialogue – needed for localization
Printmaster Approved final master mix
Dubbing guide Speaker timing, tone, pronunciation, pauses
Loudness report Evidence that the final mix meets buyer/platform loudness targets

Stem Extraction Implementation

The Post-Processor component extracts these elements based on the Audio Planning Template. Here is a simplified representation of the extraction logic:

def extract_stems(generated_clip, template):
    stems = {
        "dialogue": clip.extract_dialogue(),
        "music": clip.extract_music(),
        "effects": clip.extract_effects(),
    }
    if "M&E" in template:
        stems["M&E"] = mix(stems["music"], stems["effects"])
    generate_loudness_report(generated_clip)
    return stems

The real implementation lives in the Seedance 2.0 CI pipelines, coordinating with the Skill OS registered in agents/openai.yaml (lines 56-64).

Quality Control and Delivery Standards

Before delivery, you must run a validation checklist defined in references/audio-post-delivery.md (lines 34-43).

Pre-Delivery Sync and Loudness Checks

Verify the following before sign-off:

  • Lip-sync against the final picture
  • Music/SFX remain in sync after frame-rate conversion
  • Dialogue is not masked by music or effects
  • No unintended source-video audio remains (especially for R2V pipelines)
  • Loudness true-peak and integrated- LUFS against the buyer's spec
  • M&E or stems exist when localization is required

Common Repairs

If checks fail, references/audio-post-delivery.md (lines 46-54) specifies standard fixes:

Failure Repair
Dialogue desync Shorten line, lock framing, limit head motion, use single speaker
Wrong speaker Tag speaker, split turns
Music overwhelms line Remove music during dialogue, keep room tone
Audio reference ignored Map @Audio1 to tempo/mood and bind a visible event
Video + audio conflict Mute video reference or restrict video to camera only
Localization impossible Plan M&E/stems and text-less picture before final edit

System Architecture

The process is coordinated by four architectural layers managed through the Seedance 2.0 Skill OS:

  1. Prompt Engine – Parses the Prompt-Level Audio sections from references/audio-guide.md and injects them into the model's conditioning via skills/seedance-audio/SKILL.md
  2. Generation Model – Produces a synchronized video-audio clip in a single pass based on the locked timing cues
  3. Post-Processor – Extracts stems, runs the Audio Planning Template, and produces the Post-Level Deliverables as specified in references/audio-post-delivery.md
  4. QC Layer – Applies the Sync & Loudness Checks and triggers Common Repairs before final sign-off

All references are linked throughout the repository via [ref:audio-post-delivery] tags, with the master configuration residing in agents/openai.yaml.

Summary

  • Prompt-Level Audio: Describe dialogue, ambience, SFX, music, and sync cues explicitly in prompts to lock timing during single-pass generation
  • Audio Planning Template: Include the standardized one-line template in every prompt to guide post-processing
  • Required Deliverables: Extract full mix, dialogue stem, music stem, effects stem, M&E, printmaster, dubbing guide, and loudness report
  • Quality Control: Verify lip-sync, loudness (true-peak and integrated- LUFS), and stem completeness before delivery
  • System Integration: The workflow is managed by the Skill OS, with specifications in references/audio-post-delivery.md and agents/openai.yaml

Frequently Asked Questions

What is the Audio Planning Template in Seedance 2.0?

The Audio Planning Template is a standardized one-line syntax that must be included in every prompt to guide the post-delivery process. It specifies dialogue details, ambience beds, SFX sync cues, music tempo, audio references, and required post deliverables like stems and loudness reports. This template is parsed by the Prompt Engine and referenced by the Post-Processor during stem extraction.

How does Seedance 2.0 handle audio synchronization during generation?

Unlike traditional pipelines that add audio after video rendering, Seedance 2.0 generates video and audio simultaneously in a single pass. The model uses explicit timing cues from the prompt (such as "door slam at 3 s") as the master clock, locking both visual and audio elements to the same temporal framework during synthesis.

What is an M&E stem and why is it required?

The M&E (Music and Effects) stem contains all audio except the original dialogue, combining only the music and effects layers. This deliverable is mandatory for localization workflows because it allows dubbing houses to replace the dialogue with translated versions while preserving the original sound design and musical score.

Where are the audio post delivery specifications defined in the repository?

The master specifications reside in references/audio-post-delivery.md, which defines the required deliverables (lines 19-27), the Audio Planning Template syntax (lines 30-33), the sync and loudness checklist (lines 34-43), and common repair procedures (lines 46-54). These specifications are integrated into the Skill OS through skills/seedance-audio/SKILL.md and registered in agents/openai.yaml (lines 56-64).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →