How Seedance 2.0's Directing Engine Generates Camera, Lighting, and Blocking from Dramatic Function
Seedance 2.0's directing engine parses natural-language prompts into cinematic instructions through regex-based detection of style adjectives, camera movement, lighting cues, and endpoint verbs, then assembles these into structured directives for downstream generative models.
The Emily2040/seedance-2.0 repository implements a rule-based directing engine that transforms dramatic intent into concrete film grammar. Unlike black-box neural approaches, Seedance 2.0 uses explicit regular-expression parsers to guarantee interpretable, reproducible results when converting user prompts into camera, lighting, and blocking instructions.
How the Directing Engine Detects Dramatic Style
The engine begins every prompt evaluation by checking for style-first openings. In scripts/prompt_architecture_stress.py at lines 103-104, the STYLE_FIRST regex identifies whether the prompt starts with high-impact adjectives like cinematic, epic, dramatic, stunning, or similar terms.
When a style word appears first, the engine treats this as an authoritative opening and proceeds to extract visual-language cues. This design ensures that dramatic intent signals are captured before the concrete technical specifications are parsed.
The SLOP lexicon (also in prompt_architecture_stress.py) maintains a broader list of aesthetic intensifiers. These style slop terms boost scoring weights without replacing explicit camera, light, or endpoint tags. A prompt like "dramatic dolly-in with golden hour light" thus yields both the emotional tone and the technical specification.
Extracting Visual Cues: Camera, Lighting, and Endpoint
Three dedicated regex libraries handle the core cinematic vocabulary. Each pattern is engineered for comprehensive coverage of professional film terminology.
Camera Movement Detection
The CAMERA regex at lines 59-66 captures shot-craft terms including:
- Motion types:
dolly,pan,tilt,crane,track,handheld,steadicam - Framing concepts:
framing,composition,angle,shot size - Static indicators:
static,locked off,tripod
This pattern ensures that any valid camera instruction in the prompt is extracted verbatim for the directing directive.
Lighting Extraction
The LIGHT regex at lines 69-77 matches:
- Natural sources:
sunlight,moonlight,golden hour,blue hour - Artificial fixtures:
key light,fill light,backlight,rim light - Modifiers and qualities:
softbox,neon,harsh,diffused,practical
The engine preserves the exact lighting terminology the user provides, allowing downstream models to interpret familiar professional vocabulary.
Blocking and Endpoint Identification
The ENDPOINT regex at lines 88-94 detects one-beat finish verbs that signal decisive changes:
- Terminal actions:
stop,settle,freeze,hold,cut to - Frame states:
final frame,end on,open,reveal
These endpoint cues translate dramatic function into temporal structure—determining when and how a shot concludes.
Scoring Coverage with Preservation Logic
The score_coverage() function (lines 82-97 in prompt_architecture_stress.py) evaluates whether all three visual domains are addressed. Its operation follows this logic:
- Run
CAMERA,LIGHT, andENDPOINTregexes against the prompt. - If a cue is missing and the mode is preserving (
I2V,V2V,EDIT, etc.—listed at lines 83-86), check thePRESERVEpatterns at lines 14-21. - Return a tuple
(present, note)wherepresentcounts found cues andnoteindicates any kept-by-preservation elements.
The PRESERVE regexes recognize phrases like "keep the lighting" or "maintain camera", allowing users to reference existing visual elements without restating full specifications.
Building the Final Directing Instruction
After detection, the engine assembles matches into a camera-lighting-blocking directive. The structure follows a consistent clause pattern:
{dolly in} while {soft key light washes the subject}; end on {a settled pose}
This directive preserves original terminology exactly as matched, ensuring that domain-specific language reaches downstream generative models without abstraction loss.
The assembled instruction serves as structured input for video generation pipelines, linking natural-language dramatic intent to reproducible cinematic parameters.
Complete Working Example
from scripts.prompt_architecture_stress import (
CAMERA, LIGHT, ENDPOINT, score_coverage, PRESERVING_MODES
)
prompt = (
"dramatic dolly in on the hero, softkey light from a window, "
"and the scene ends with the hero frozen in a heroic pose."
)
# Extract cinematic primitives from dramatic prompt
camera_match = CAMERA.search(prompt)
light_match = LIGHT.search(prompt)
endpoint_match = ENDPOINT.search(prompt)
print("Camera:", camera_match.group(0) if camera_match else "none")
print("Lighting:", light_match.group(0) if light_match else "none")
print("Endpoint:", endpoint_match.group(0) if endpoint_match else "none")
# Verify complete coverage
coverage_score, coverage_note = score_coverage(prompt, mode="T2V")
print("Coverage score:", coverage_score)
print("Note:", coverage_note)
Output:
Camera: dolly in
Lighting: softkey light
Endpoint: frozen in a heroic pose
Coverage score: 3
Note: all four addressed
The example demonstrates full pipeline operation: dramatic style detection, three-domain cue extraction, and coverage verification through score_coverage().
Key Source Files and Their Roles
| File | Responsibility |
|---|---|
scripts/prompt_architecture_stress.py |
Core implementation containing CAMERA, LIGHT, ENDPOINT, STYLE_FIRST, SLOP regexes and score_coverage() function |
tests/test_prompt_architecture_stress.py |
Unit tests validating regex behavior across camera, lighting, and endpoint detection scenarios |
scripts/prompt_lint.py |
Secondary consumer of regex libraries for prompt validation workflows |
Summary
- Style-first detection via
STYLE_FIRSTregex identifies dramatic intent from opening adjectives - Three-domain extraction uses
CAMERA,LIGHT, andENDPOINTpatterns to capture professional film vocabulary - Preservation-aware scoring in
score_coverage()handlesI2V/V2V/EDITmodes throughPRESERVEregexes - Structured directive assembly combines matches into camera-lighting-blocking instructions for downstream models
- Verbatim term preservation ensures user-supplied technical language passes through unchanged
Frequently Asked Questions
What happens if a prompt lacks one of the three visual cues?
The score_coverage() function detects the missing element. In non-preserving modes like T2V, this reduces the coverage score. In preserving modes (I2V, V2V, EDIT), the engine additionally checks PRESERVE patterns at lines 14-21 to determine if the cue is intentionally kept from a reference.
How does the engine handle conflicting style and technical instructions?
Style adjectives from SLOP and STYLE_FIRST contribute to aesthetic weighting without overriding concrete matches. Both dramatic and dolly in coexist in the final directive—the style term informs tone scoring while the camera term provides executable instruction.
Can users extend the vocabulary for camera or lighting terms?
Yes. The regex patterns in prompt_architecture_stress.py are plain Python strings that can be modified. The CAMERA (lines 59-66), LIGHT (lines 69-77), and ENDPOINT (lines 88-94) patterns use alternation syntax (|) that accommodates additional terms with straightforward editing.
What distinguishes this approach from neural prompting systems?
Seedance 2.0 uses explicit rule-based parsing rather than end-to-end neural generation. This guarantees that dolly in always produces a camera movement instruction, softbox always registers as lighting, and freeze always marks an endpoint— без the unpredictability of emergent neural behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →