How to Use the Seedance 2.0 Capability Map for Prompt Design: A Technical Guide
The Seedance 2.0 capability map is a four-layer verification framework in references/capability-map.md that separates what the model can do from what any single surface exposes, enabling precise prompt design that exploits official strengths while working around documented limits.
The Capability Map is the central reference document for the Emily2040/seedance-2.0 repository. It defines what the Seedance 2.0 video generation model can do, how to access those capabilities through specific syntax, and what limitations must shape your prompt architecture. Unlike provider documentation that makes universal claims, the map forces you to verify capabilities across four independent layers before committing to a design strategy.
The Four Verification Layers
According to the source code in references/capability-map.md lines 11-14, every capability must pass through four verification stages:
| Layer | Question Answered | Evidence Required |
|---|---|---|
| Model capability | Can the model do this at all? | Model-level documentation describing the modality |
| Surface access | Does this specific surface expose it? | Availability per operation, not guaranteed across siblings |
| Request syntax | How is the capability expressed? | Example showing the authoring form |
| Returned adherence | Did the output actually follow it? | Observed take as evidence, not guarantee |
This structure prevents the common failure mode of assuming a feature works everywhere just because it appears in official documentation.
Designing Into Model Strengths: Extraction Moves
The capability map enumerates concrete extraction moves you should preload before drafting any prompt. These are located in references/capability-map.md starting at line 30.
Multi-Shot Generation
Shot 1:/2:/3: labels · one action · one camera · standard tier [field] · 10-15 s / auto · shots × seconds budget
This extraction move from line 34 enables multiple shots in a single API call—a core efficiency feature of Seedance 2.0.
# Prompt
Shot 1: A close-up of a ceramic teapot steaming, the camera dollies forward, soft piano notes rise.
Shot 2: The teapot tilts, pouring steaming tea into a glass, the camera tracks left, a faint chime echoes.
Shot 3: A hand lifts the glass, sunlight glints on the surface, the camera pans up to reveal a sunrise.
duration: auto
Why this works:
Shot 1:/2:/3:labels satisfy the Request syntax layer- One primary action + one camera move per shot follows the Model capability extraction move
duration: autolets the model allocate roughly 4-6 seconds per shot, respecting the budget guideline
The exact syntax for these labels is defined in references/multishot-grammar.md line 5, which the capability map references as the authoritative request format.
Native Synced Audio
Name specific sounds, keep dialogue short, prioritize SFX > music > dialogue
From line 35 of the capability map, this move ensures audio elements render correctly:
# Prompt
Shot 1: A city street at dusk, a single car horn blares, the camera pans right to follow the car.
Audio: "car horn" – crisp, short, on-screen.
Shot 2: Neon signs flicker, a soft synth pad rises, the camera tilts up.
Audio: "synth pad" – sustained, low-key.
duration: auto
Explicit Audio: tags implement the extraction move. The capability map emphasizes keeping dialogue minimal and ranking sound types by reliability.
First/Last Frame Locking
Lock endpoints, prompt → travel → resolve; use transformations & match-cuts
From line 39, this technique constrains generation between fixed visual states:
# Prompt
Start frame: a handheld shot of a wooden door, light spilling through the crack.
Shot 1: The door opens slowly, camera tracks forward, a creak is heard.
Shot 2: Inside, a candle flickers, the camera tilts down, soft wind whistles.
End frame: the candle blows out, darkness settles, the camera pulls back.
duration: 12s
The First/last frame extraction move requires explicit Start frame: and End frame: declarations. Each intermediate shot maintains single-action/single-camera discipline per line 40's Literal camera verbs rule.
Physics-Based Description
Use physical verbs & consequences, not pose adjectives
Line 41 of the capability map notes that describing what happens (verbs + physical consequences) outperforms static pose descriptions for motion coherence.
Style-Specific Modes
| Mode | Extraction Move | Source |
|---|---|---|
| 2D/anime | Cel-style grammar, no lens/DOF talk | Line 44 |
| Multilingual | Anchor texture/mood in Chinese, keep reference tags exact | Line 46 |
Designing Around Model Limits
When chaining generations or pushing duration boundaries, the capability map at line 50 recommends specific mitigations:
- Limit clip length to avoid continuity drift across generated segments
- Preserve exact reference roles (e.g., "motion only, no appearance")
- Re-anchor on schedule at the scene's chain-depth cap
- Treat surface-specific duration caps as facts of that surface, not of the underlying model
These constraints reflect the Returned adherence layer—what you observe in actual takes, not what documentation promises.
Integration With the Design Pipeline
The capability map sits at the foundation of Seedance 2.0's documented workflow. SKILL.md line 50 shows the "Capability check" step explicitly loading the map before any shot planning occurs. README.md line 69 positions it as the first of three design aids alongside the fidelity-allocation model and model-mechanics reference.
The complete pipeline runs:
- Consult the Capability Map → identify extraction moves
- Allocate fidelity (
references/allocation-model.md) → distribute rendering budget - Apply concrete grammar (
references/multishot-grammar.md) → implement syntax - Generate → execute API call
- Validate adherence → use
prompt-lintand eval-run tools
Summary
- The capability map (
references/capability-map.md) provides a four-layer verification framework that separates model capabilities from surface-specific availability - Extraction moves like multi-shot labels, audio tags, and endpoint locking translate capabilities into concrete prompt syntax
- The request syntax layer depends on
references/multishot-grammar.mdfor exact formatting rules - Design-around strategies for chaining and duration limits are based on observed behavior, not documentation claims
- The map is explicitly loaded during the "Capability check" phase documented in
SKILL.md
Frequently Asked Questions
What is the difference between "official" and "[field]" markers in the capability map?
Official capabilities are documented by the model provider and generally available across surfaces. [field] markers indicate capabilities verified through production use but potentially subject to surface-specific restrictions. According to references/capability-map.md, you should verify [field] capabilities through the Surface access layer before depending on them.
Can I use multi-shot prompts on any Seedance 2.0 API surface?
Not necessarily. The capability map's Surface access layer requires explicit verification. While the Model capability layer confirms multi-shot generation exists, individual surfaces may not expose it. Check references/api-status.md for surface-specific resolution and feature availability before designing your prompt architecture.
Why does the capability map distinguish "one take is evidence, not a guarantee"?
This reflects the Returned adherence verification layer. A single successful generation proves a capability can work under specific conditions, but stochastic variation means subsequent runs may differ. For production reliability, you should accumulate multiple observed takes or implement the design-around strategies documented at line 50 of the capability map.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →