How Seedance 2.0 Model Mechanics Explain Generation Rules: A Technical Guide
Seedance 2.0 generation rules emerge from eight core diffusion model mechanisms defined in references/model-mechanics.md, each governing how conditioning resources allocate across tokens, references, and time.
The Emily2040/seedance-2.0 repository treats generation rules not as arbitrary constraints but as observable properties of the underlying diffusion architecture. Understanding these mechanisms lets prompt engineers diagnose failures and optimize outputs without trial-and-error guessing.
The Eight Mechanisms That Drive Every Generation
The model-mechanics.md reference document structures all generation behavior around eight mental models. Each mechanism translates directly into actionable prompt-engineering rules.
1. Attention Is a Budget
Diffusion conditioning operates as a finite resource competition. Early tokens capture disproportionate influence; filler words ("slop") consume capacity without improving output.
Practical rules derived:
- Lead with subject and action
- Compress prompts to information-dense phrasing
- Strip adjectives that don't alter visible pixels
# Correct: subject/action first, no wasted tokens
prompt = "A heroic astronaut stepping onto a moon crater, photorealistic, soft rim light."
# Avoid: "Here is a scene where we see an astronaut who is..."
2. Generation Pulls Toward the Familiar
The model samples preferentially near high-density training clusters. Common visual combinations cost less compute and stabilize better; rare intersections create instability and flicker.
Practical rules derived:
- Invoke established style vocabularies (film noir, cel-animation, golden hour)
- Lock style with identical anchor phrases per shot
- Budget extra validation for novel concept combinations
anchor = "cinematic 8K, film noir lighting, high contrast"
prompt = f"{anchor}, a lone detective in a rain-soaked alley, neon signs, wet pavement."
# Reuse `anchor` unchanged across every shot in the sequence
3. There Is No NOT
Negation operators still activate their target concepts at the embedding level. The model cannot "subtract" what it has already encoded.
Practical rules derived:
- Describe only present elements
- Route exclusions to platform-level constraints (
no watermark,no on-screen text)
# WRONG: still activates "blood" and "wounded" strongly
bad_prompt = "A soldier with no blood visible, not wounded"
# GOOD: controls composition through positive description
prompt = "A wounded soldier with a bandaged arm, blood spatter on the ground."
4. Time Is a Trajectory Prior
Motion generation assumes smooth, physically coherent cause-and-effect progressions. Discontinuous instructions lack the temporal gradient the sampler expects.
Practical rules derived:
- Structure prompts around single physical causes with visible consequences
- Replace "jump-cut" descriptions with burst versus held timing grammar
# Cause flows to effect along continuous path
prompt = "A marble rolls down a marble-styled staircase, the camera follows its path, slow motion."
5. Errors Compound
Each frame conditions autoregressively on neighbors. Identity drift accumulates across clips; feedback loops amplify deviations when outputs re-enter the conditioning chain.
Practical rules derived:
- Re-anchor long clips to the original reference, not successive outputs
- Hard-limit generation chains to ≤4–5 iterations before reference reset
reference_image = "ref_hero.png"
# Re-anchor to original, don't chain from previous generation
prompt = f"[ref:{reference_image}] Hero portrait, keep original face, new background sunrise."
6. References Outrank Text Where They Overlap
Image, video, and audio references supply dense conditioning signals that override textual descriptions at conflict points. Redundant text creates interference rather than reinforcement.
Practical rules derived:
- Prompt only what references cannot encode (temporal change, sound design, constraints)
- Explicitly flag what must not transfer from reference assets
# Don't re-describe the cityscape; add only new elements
prompt = "[ref:cityscape.jpg] Night city skyline, add a hovering drone, keep building silhouettes unchanged."
7. Detail Capacity Scales With Screen Area
Representational allocation is spatially proportioned. Small regions receive reduced capacity; fine detail degrades first under motion stress.
Practical rules derived:
- Scale hero subjects to dominate frame area
- Isolate tiny critical details in dedicated shots or zoom levels
# Dedicated shot preserves watch detail that would blur in wide frame
prompt = "Close-up of a vintage watch face, 2-second hold, then cut to wide shot of the explorer's hand."
8. Audio and Video Are Generated Together
Sound synthesis shares the diffusion trajectory with pixels. Named audio events function as temporal anchors; dialogue requires facial stability and brevity constraints.
Practical rules derived:
- Name each shot's soundscape explicitly
- Use audio timing as editorial clock; constrain spoken lines
prompt = """Shot 1: A thunderstorm, heavy rain, and distant thunder rumble.
Shot 2: Quiet library, pages flipping softly."""
Mechanism-Indexed Diagnosis
When generation outputs deviate from intent, model-mechanics.md provides a diagnostic lookup table mapping symptom to dominant mechanism and corrective lever:
| Symptom | Dominant Mechanism | Lever to Apply |
|---|---|---|
| Subject ignored, style dominant | Attention budget | Reorder: subject first, style last |
| Style drift across sequence | Pull toward familiar | Repeat identical anchor phrase |
| Unwanted element appearing | No NOT | Remove negation, describe positively |
| Discontinuous motion | Trajectory prior | Unify around single cause-effect |
| Identity drift across clip | Errors compound | Re-anchor to original reference |
| Reference aesthetic overriding text | Reference outranks text | Strip overlapping descriptions |
| Detail loss in small regions | Detail capacity | Isolate in dedicated shot |
| Audio-visual sync failure | Audio-video joint | Name sound events, shorten dialogue |
Repository Implementation
The mechanics document directly shapes automated validation in three linked components:
scripts/generation_run_check.py— CLI validator enforcing mechanics constraints (propergeneration_mode, slop detection, clip length limits)schemas/generation-run.schema.json— JSON Schema encoding required fields derived from the eight mechanismstests/test_generation_run_check.py— Unit tests asserting correct violation flagging
# Example: validation script checking attention budget compliance
# From scripts/generation_run_check.py
def check_slop_density(prompt: str) -> bool:
"""Flag prompts with filler words that waste attention budget."""
slop_patterns = ["here is", "scene where", "we see", "image of"]
return not any(pattern in prompt.lower() for pattern in slop_patterns)
Summary
- Seedance 2.0 generation rules originate from eight mechanisms in
references/model-mechanics.md, not arbitrary policy - Attention mechanisms demand dense, front-loaded prompts with subject-first ordering
- Trajectory priors require cause-effect coherence; error compounding hard-limits generation chains to 4–5 clips
- Reference conditioning overrides text at overlap points; spatial capacity scales with screen area
- Audio-video joint generation makes sound design a first-class timeline control mechanism
- Validation tools
generation_run_check.pyandgeneration-run.schema.jsonencode these mechanics into enforceable constraints
Frequently Asked Questions
How do I fix style drift across multiple shots in Seedance 2.0?
Apply mechanism #2 (pull toward familiar) by using an identical anchor phrase in every shot. The model stabilizes output by sampling near repeated conditioning clusters. Lock style with a fixed prefix like cinematic 8K, film noir lighting, high contrast before shot-specific descriptions.
Why does my negative prompt ("no watermark") still produce watermarks?
Mechanism #3 (there is no NOT) explains this: negation activates concepts rather than suppressing them. The embedding for "watermark" enters conditioning regardless of negation syntax. Use platform-level constraints in generation-run.schema.json fields, or describe the clean desired state positively.
When should I use reference images versus text descriptions?
Follow mechanism #6 (reference outranks text). References dominate where they overlap with text, so prompt only what the reference cannot convey: temporal dynamics, sound events, or constraints on what not to transfer. In model-mechanics.md terms, references are "dense conditioning" that overshadows textual "sparse conditioning" at conflict points.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →