How Animation Payoff Timing and Narration Sync Work in video-use

video-use synchronizes overlay animations to spoken narration by starting animations reveal_duration seconds before a payoff word timestamp, ensuring the final frame lands exactly when the key phrase is spoken.

The browser-use/video-use repository treats narration as the master clock for all visual overlays. This approach ensures that animated diagrams, text cards, and complex visuals reinforce the spoken narrative at precisely the right moment. Understanding how animation payoff timing and narration sync interact is essential for producing professional-grade educational content where visuals and audio feel seamlessly integrated.

The Two Core Synchronization Rules

The synchronization strategy defined in SKILL.md relies on two complementary timing constraints that govern how overlays appear relative to the spoken track.

Sync-to-Narration Timing

Every narration-driven overlay must persist long enough for viewers to comfortably read and comprehend the content. According to the guidelines in SKILL.md at line 218, durations follow these floors:

  • 3 seconds – absolute minimum for any readable overlay
  • 5–7 seconds – typical duration for simple text cards
  • 8–14 seconds – required for complex diagrams with multiple elements

This rule ensures accessibility and prevents cognitive overload, regardless of when the animation technically begins.

Animation Payoff Timing

The payoff refers to the final frame of an animation—the moment when the complete diagram or text block is fully revealed. To maximize impact, this payoff frame must coincide exactly with the spoken payoff word (the key term or concept being explained).

As implemented in SKILL.md at line 224, the overlay start time is calculated by subtracting the animation's duration from the payoff word's timestamp:

start_in_output = payoff_timestamp - reveal_duration

This guarantees that after reveal_duration seconds of animation (e.g., drawing, fading, or building elements), the viewer sees the completed visual exactly when the narrator utters the critical phrase.

Implementation in the Rendering Pipeline

The actual timestamp alignment happens in helpers/render.py, which processes the Edit Decision List (EDL) and constructs the final FFmpeg command.

PTS Shifting for Frame Alignment

When helpers/render.py processes overlay segments (lines 7–9), it applies a PTS shift using the setpts filter:

setpts=PTS-STARTPTS+T/TB

This formula adjusts the Presentation Timestamp so that frame 0 of the overlay video aligns with the calculated start_in_output value in the EDL timeline. The overlay essentially becomes a contiguous stream that starts playing at the precise moment needed to hit the payoff word.

Filter Graph Ordering

The rendering pipeline applies the subtitle filter last to ensure that captions remain visible above any animated overlays. This prevents animation elements from obscuring text that viewers need to read.

Calculating Overlay Start Times in Practice

To implement payoff timing in your own EDL workflows, use the following pattern when building overlay entries. This Python snippet demonstrates how to compute the correct start_in_output value given a payoff word timestamp and desired reveal duration:

import json
from pathlib import Path

# --------------------------------------------------------------

# Helper: compute overlay start based on payoff word timestamp

# --------------------------------------------------------------

def make_overlay(entry, payoff_ts: float, reveal_dur: float) -> dict:
    """
    entry: dict from the EDL's `overlays` list (may contain file path, duration)
    payoff_ts: timestamp of the spoken payoff word (seconds)
    reveal_dur: how many seconds before the payoff the animation should start
    """
    start = max(0.0, payoff_ts - reveal_dur)          # never start before 0

    entry["start_in_output"] = round(start, 3)       # precision matches ffprobe output

    entry["duration"] = round(reveal_dur + 0.5, 3)   # include a short tail after payoff

    return entry

# --------------------------------------------------------------

# Example usage – building an EDL for a simple card overlay

# --------------------------------------------------------------

edl_path = Path("edit/edl.json")
edl = json.loads(edl_path.read_text())

# Suppose we have an overlay video already rendered at `edit/animations/slot_3/render.mp4`

overlay = {
    "file": "edit/animations/slot_3/render.mp4",
    "duration": 0,           # placeholder – will be overwritten

    "start_in_output": 0,    # placeholder – will be overwritten

}

# Payoff word occurs at 12.34 s in the narration; we want a 2-second reveal

payoff_timestamp = 12.34
reveal_duration = 2.0

edl["overlays"].append(make_overlay(overlay, payoff_timestamp, reveal_duration))

# Save the modified EDL

edl_path.write_text(json.dumps(edl, indent=2))
print("EDL updated with payoff-timed overlay.")

When helpers/render.py processes this EDL, it shifts the overlay's PTS so the animation begins at 10.34 seconds (12.34 - 2.0), revealing for exactly 2 seconds until the payoff word at 12.34 seconds. The optional 0.5-second tail keeps the final frame visible briefly after the narrator finishes the key phrase.

Summary

  • Animation payoff timing in video-use requires calculating start_in_output = payoff_timestamp - reveal_duration to align final animation frames with spoken keywords.
  • Sync-to-narration rules mandate minimum durations of 3 seconds (floor), with 5–7 seconds for simple content and 8–14 seconds for complex diagrams per SKILL.md.
  • The rendering pipeline in helpers/render.py uses FFmpeg's setpts=PTS-STARTPTS+T/TB to align overlay frame 0 with EDL start times.
  • Subtitle filters are applied last in the filter chain to prevent overlays from hiding captions.
  • EDL files stored at edit/edl.json serve as the bridge between narrative timestamps and precise animation timing.

Frequently Asked Questions

What is a "payoff word" in the context of video-use?

A payoff word is the specific term or concept in the narration that represents the key learning moment for a given visual. In video-use, animations are timed so that their final "payoff" frame—showing the completed diagram or text—appears exactly when the narrator speaks this word, creating immediate visual-audio reinforcement.

How long should narration-driven overlays remain on screen?

According to SKILL.md, overlays must last at least 3 seconds to be readable, with 5–7 seconds being typical for simple cards and 8–14 seconds for complex diagrams. These durations are independent of animation length; if an animation takes 2 seconds to reveal, the total overlay time must still meet the minimum readability threshold.

Why does the subtitle filter run last in the rendering pipeline?

The rendering pipeline in helpers/render.py applies the subtitle filter after all video and overlay filters to ensure that captions remain visible on top of any animated content. This ordering prevents complex diagrams or visual effects from obscuring text that viewers need to read.

How do I adjust reveal duration for complex Manim animations?

For Manim-based diagrams (referenced in skills/manim-video/SKILL.md), calculate the reveal_duration to match the actual render time of the build sequence, then add a small buffer (0.5–1.0 seconds) to the total overlay duration. This ensures the diagram finishes constructing exactly at the payoff word while remaining visible long enough for the viewer to study the final result.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →