Seedance 2.0 Audio Post Delivery Process: From Prompt to Broadcast Master
The Seedance 2.0 audio post delivery process generates synchronized video and audio in a single pass using explicit prompt-level audio descriptions, then extracts professional stems and metadata according to specifications defined in references/audio-post-delivery.md to ensure broadcast-ready output.
Seedance 2.0 is an open-source generative media framework that treats audio as a first-class element of every prompt. According to the Emily2040/seedance-2.0 repository, the workflow splits generation and delivery into two distinct layers: prompt-level audio conditioning that locks timing during synthesis, and post-level deliverable extraction that prepares assets for downstream editors and automated pipelines.
Prompt-Level Audio Conditioning
Seedance 2.0 requires explicit audio descriptions in prompts to synchronize generation. When you write a prompt, you describe the type of audio needed so the model can lock timing and mood during the single-pass generation process.
Audio Categories and Prompt Syntax
The model recognizes five distinct audio categories that must be described explicitly in prompts, as documented in references/audio-guide.md (lines 11-14):
- Dialogue: Short quoted lines, one speaker per turn
- Ambience: Room tone, traffic, rain, crowd bed, machine hum
- SFX: One-or-two concrete sounds tied to a visible action
- Music: Broad tempo/energy/instrumentation, or explicit "silence"
- Sync cue: Visible event that marks a beat (e.g., "door slam at endpoint")
Example prompt structure:
Dialogue: "I'm late!" Sound: low room tone + distant rain. SFX: door slam at 3 s. Music: none until after the line.
The model generates video and audio together in a single pass, so the timing specified in the prompt becomes the "clock" for the entire clip.
The Audio Planning Template
Every prompt should include a one-line Audio Planning Template that guides the post-process. According to references/audio-post-delivery.md (lines 30-33), the template follows this syntax:
Dialogue: [speaker/line/language]. Ambience: [bed]. SFX: [visible sync cue].
Music: [tempo/mood or none]. Reference: [Audio1 role].
Post: [stems/M&E/loudness/sync notes].
When this template is parsed by the Prompt Engine (implemented in skills/seedance-audio/SKILL.md), it injects audio conditioning into the model's generation parameters.
Post-Level Audio Deliverables
After generation, you must hand off a set of stems and metadata so downstream editors can finish the project for broadcast or platform delivery.
Required Stems and Metadata
The specific deliverables are listed in references/audio-post-delivery.md (lines 19-27):
| Deliverable | Purpose |
|---|---|
| Full mix | Final stereo/5.1/Atmos (or platform-required) mix |
| Dialogue stem | Isolated spoken lines for dubbing, edits, cleanup |
| Music stem | Music-only layer |
| Effects stem | All SFX and designed sound |
| M&E | Music + Effects without original dialogue – needed for localization |
| Printmaster | Approved final master mix |
| Dubbing guide | Speaker timing, tone, pronunciation, pauses |
| Loudness report | Evidence that the final mix meets buyer/platform loudness targets |
Stem Extraction Implementation
The Post-Processor component extracts these elements based on the Audio Planning Template. Here is a simplified representation of the extraction logic:
def extract_stems(generated_clip, template):
stems = {
"dialogue": clip.extract_dialogue(),
"music": clip.extract_music(),
"effects": clip.extract_effects(),
}
if "M&E" in template:
stems["M&E"] = mix(stems["music"], stems["effects"])
generate_loudness_report(generated_clip)
return stems
The real implementation lives in the Seedance 2.0 CI pipelines, coordinating with the Skill OS registered in agents/openai.yaml (lines 56-64).
Quality Control and Delivery Standards
Before delivery, you must run a validation checklist defined in references/audio-post-delivery.md (lines 34-43).
Pre-Delivery Sync and Loudness Checks
Verify the following before sign-off:
- Lip-sync against the final picture
- Music/SFX remain in sync after frame-rate conversion
- Dialogue is not masked by music or effects
- No unintended source-video audio remains (especially for R2V pipelines)
- Loudness true-peak and integrated- LUFS against the buyer's spec
- M&E or stems exist when localization is required
Common Repairs
If checks fail, references/audio-post-delivery.md (lines 46-54) specifies standard fixes:
| Failure | Repair |
|---|---|
| Dialogue desync | Shorten line, lock framing, limit head motion, use single speaker |
| Wrong speaker | Tag speaker, split turns |
| Music overwhelms line | Remove music during dialogue, keep room tone |
| Audio reference ignored | Map @Audio1 to tempo/mood and bind a visible event |
| Video + audio conflict | Mute video reference or restrict video to camera only |
| Localization impossible | Plan M&E/stems and text-less picture before final edit |
System Architecture
The process is coordinated by four architectural layers managed through the Seedance 2.0 Skill OS:
- Prompt Engine – Parses the Prompt-Level Audio sections from
references/audio-guide.mdand injects them into the model's conditioning viaskills/seedance-audio/SKILL.md - Generation Model – Produces a synchronized video-audio clip in a single pass based on the locked timing cues
- Post-Processor – Extracts stems, runs the Audio Planning Template, and produces the Post-Level Deliverables as specified in
references/audio-post-delivery.md - QC Layer – Applies the Sync & Loudness Checks and triggers Common Repairs before final sign-off
All references are linked throughout the repository via [ref:audio-post-delivery] tags, with the master configuration residing in agents/openai.yaml.
Summary
- Prompt-Level Audio: Describe dialogue, ambience, SFX, music, and sync cues explicitly in prompts to lock timing during single-pass generation
- Audio Planning Template: Include the standardized one-line template in every prompt to guide post-processing
- Required Deliverables: Extract full mix, dialogue stem, music stem, effects stem, M&E, printmaster, dubbing guide, and loudness report
- Quality Control: Verify lip-sync, loudness (true-peak and integrated- LUFS), and stem completeness before delivery
- System Integration: The workflow is managed by the Skill OS, with specifications in
references/audio-post-delivery.mdandagents/openai.yaml
Frequently Asked Questions
What is the Audio Planning Template in Seedance 2.0?
The Audio Planning Template is a standardized one-line syntax that must be included in every prompt to guide the post-delivery process. It specifies dialogue details, ambience beds, SFX sync cues, music tempo, audio references, and required post deliverables like stems and loudness reports. This template is parsed by the Prompt Engine and referenced by the Post-Processor during stem extraction.
How does Seedance 2.0 handle audio synchronization during generation?
Unlike traditional pipelines that add audio after video rendering, Seedance 2.0 generates video and audio simultaneously in a single pass. The model uses explicit timing cues from the prompt (such as "door slam at 3 s") as the master clock, locking both visual and audio elements to the same temporal framework during synthesis.
What is an M&E stem and why is it required?
The M&E (Music and Effects) stem contains all audio except the original dialogue, combining only the music and effects layers. This deliverable is mandatory for localization workflows because it allows dubbing houses to replace the dialogue with translated versions while preserving the original sound design and musical score.
Where are the audio post delivery specifications defined in the repository?
The master specifications reside in references/audio-post-delivery.md, which defines the required deliverables (lines 19-27), the Audio Planning Template syntax (lines 30-33), the sync and loudness checklist (lines 34-43), and common repair procedures (lines 46-54). These specifications are integrated into the Skill OS through skills/seedance-audio/SKILL.md and registered in agents/openai.yaml (lines 56-64).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →