The 12 Hard Production Rules in video-use: Why They Matter for Broadcast-Ready Output
TLDR: The video-use repository enforces 12 non-negotiable hard production rules that prevent silent failures like audio pops, missing subtitles, and double re-encoding, ensuring every LLM-driven edit produces broadcast-ready video without costly manual rework.
The browser-use/video-use project defines a strict set of I2 hard production rules documented in [SKILL.md](https://github.com/browser-use/video-use/blob/main/SKILL.md#L18-L34) that serve as guardrails for its video rendering pipeline. These rules are non-negotiable constraints implemented in helpers/render.py and helpers/transcribe.py that guarantee deterministic timeline construction, lossless processing, and professional AV sync. Violating any rule results in audible artifacts, misaligned animations, or destructive re-encoding that automated quality checks cannot detect.
The 12 Hard Production Rules Explained
1. Subtitles Are Applied Last in the Filter Chain
Subtitles must be applied after every overlay in the FFmpeg filter chain. This rule prevents graphics and animations from hiding caption text, which would cause silently-failing subtitles that are invisible to the viewer.
In helpers/render.py, the pipeline constructs the filter complex so that [vid]subtitles=master.srt executes only after all overlay operations complete.
ffmpeg -i video.mp4 -i overlay.webm -filter_complex "[0:v][1:v]overlay[vid];[vid]subtitles=master.srt" -c:a copy out.mp4
2. Per-Segment Extract with Lossless Concatenation
Extract segments using -c copy and concatenate losslessly, never use a single-pass filtergraph with overlays. A single filtergraph would force a full re-encode of the entire source, doubling encoding time and degrading quality through generational loss.
The implementation in helpers/render.py first extracts each segment losslessly, then concatenates via the concat demuxer:
# Extract losslessly
ffmpeg -i src.mp4 -ss $START -to $END -c copy seg_$i.mp4
# Concat without re-encoding
ffmpeg -f concat -safe 0 -i list.txt -c copy out.mp4
3. 30ms Audio Fades at Every Segment Boundary
Apply 30ms micro-fades at all cut edges using afade=t=in:st=0:d=0.03 and afade=t=out:st={dur-0.03}:d=0.03. Without these fades, abrupt amplitude changes create audible "pop" clicks at each cut, which are unacceptable for professional video output.
This is enforced when helpers/render.py processes each segment extract:
ffmpeg -i seg.mp4 -af "afade=t=in:st=0:d=0.03,afade=t=out:st=$(duration-0.03):d=0.03" -c:v copy seg_faded.mp4
4. Overlays Use PTS Reset to Window Start
Overlay videos must use setpts=PTS-STARTPTS+T/TB to shift the overlay's frame 0 to its window start time. If the overlay timing isn't reset, the animation will start in the middle of its timeline, displaying incorrect frames during the overlay window.
ffmpeg -i overlay.webm -filter:v "setpts=PTS-STARTPTS+5/TB" shifted_overlay.webm
5. Master SRT Uses Output-Timeline Offsets
Subtitle timestamps must be calculated relative to the final output timeline, not the original source. The formula output_time = word.start - segment_start + segment_offset ensures captions remain synchronized after concatenation.
This logic is implemented in helpers/render.py when building the master SRT file, ensuring subtitles align with the edited timeline defined in edl.json.
6. Never Cut Inside a Word
All cut edges must snap to word boundaries from the Scribe transcript stored in takes_packed.md. Cutting mid-word creates jarring audio-visual glitches and destroys the precision of transcript-driven editing.
The planner consults takes_packed.md generated by helpers/pack_transcripts.py to locate nearest word timestamps before rendering.
7. Pad Every Cut Edge (30-200ms)
Apply 30-200ms of padding to every cut edge. This working window absorbs Scribe timestamp drift (approximately 50-100ms) and provides a safety margin for visual continuity between segments.
helpers/render.py automatically applies this padding when generating per-segment extracts, ensuring seamless transitions in the final concat.
8. Word-Level Verbatim ASR Only
Use only word-level verbatim ASR output, never SRT/phrase mode or normalized fillers. Phrase-level output lacks the sub-second timing precision required for accurate cuts and subtitle offsets.
The transcription helper in helpers/transcribe.py forces verbatim=true in all Scribe API requests, ensuring cached transcripts under edit/transcripts/ contain millisecond-accurate word timings.
9. Cache Transcripts Per Source
Never re-transcribe unless the source file changes. Re-transcribing is expensive (network + CPU) and can produce slightly different timestamps, breaking deterministic cuts and the edl.json consistency.
Cached JSON files live under edit/transcripts/ and are keyed to source file hashes, as managed by helpers/transcribe.py.
10. Parallel Sub-Agents for Multiple Animations
Render multiple animations in parallel using sub-agents, never sequentially. Serial rendering would multiply wall-time by the number of overlays; parallel agents keep overall latency close to the duration of the slowest animation.
helpers/render.py spawns N animation agents via the Agent tool to process overlays concurrently.
11. Strategy Confirmation Before Execution
Never execute cuts until the user approves the plain-English plan. This guarantees that the LLM's reasoning is visible and validated, preventing surprise edits and ensuring alignment with editorial intent.
The skill pauses after generating a textual plan in SKILL.md workflow, requiring explicit confirmation before touching source files.
12. Session Outputs Isolated to <videos_dir>/edit/
Write all session outputs to <videos_dir>/edit/, never inside the video-use/ project directory. This keeps source footage immutable, separates project artifacts from the repository code, and simplifies cleanup and version control.
Rendered files are stored under edit/ as specified in the directory layout documented in SKILL.md.
Architectural Rationale
These 12 hard production rules form a guardrail layer that transforms a flexible LLM-driven workflow into a production-grade system:
- Deterministic timeline: By cutting only at word boundaries, padding cuts, and using cached transcripts, the pipeline builds a reproducible edit decision list (
edl.json) - Single-pass quality: Extract-then-concat avoids double-encoding, preserving source fidelity while allowing per-segment grading and fades
- Audio-visual sync: The 30ms fade and overlay PTS shift guarantee smooth transitions and correctly timed animations, preventing "pop" and "frame jump" artifacts
- Subtitle reliability: Placing subtitles last and computing offsets from the final timeline ensures captions are always visible and perfectly aligned
Summary
- Subtitles last: Apply captions after all overlays to prevent hidden text
- Lossless concat: Extract segments with
-c copybefore concatenating to avoid generational loss - Audio fades: Mandatory 30ms fades prevent audible pops at cut points
- PTS reset: Reset overlay timestamps to align animations with their window start
- Timeline math: Calculate subtitle offsets against the final output timeline
- Word boundaries: Snap all cuts to transcript word boundaries to avoid jarring edits
- Edge padding: Add 30-200ms padding to absorb timestamp drift
- Verbatim ASR: Force word-level transcription for sub-second precision
- Transcript caching: Cache transcripts by source hash to ensure deterministic renders
- Parallel rendering: Use concurrent agents for overlay generation to minimize latency
- User confirmation: Require explicit approval of the edit plan before execution
- Directory isolation: Keep all outputs in
edit/to maintain immutable source footage
Frequently Asked Questions
What happens if I apply subtitles before overlays in video-use?
Applying subtitles before overlays violates Rule 1 and results in graphics covering the caption text. Because this failure is silent (the video renders without errors but hides the subtitles), it often requires a full manual re-render to fix. The helpers/render.py implementation explicitly orders the filter chain to prevent this.
Why does video-use require 30ms audio fades instead of using crossfades?
The 30ms micro-fades specified in Rule 3 eliminate audible "pop" clicks caused by DC offset shifts at cut boundaries. Crossfades would blend adjacent audio content, which is unnecessary for most cuts and would alter the timing. The 30ms duration is optimized to absorb Scribe timestamp drift without being perceptible to listeners.
How does video-use prevent double-encoding of source footage?
Rule 2 mandates per-segment extraction using -c copy followed by lossless concatenation. This approach preserves the original source codec and quality, whereas a single-pass filtergraph with overlays would force a full re-encode of the entire timeline. The helpers/render.py script manages this two-stage process to maintain generation-loss-free output.
Where are the hard production rules documented in the repository?
The canonical specification for all 12 rules resides in SKILL.md at lines 18-34, which defines them as "Hard Rules (production correctness — non-negotiable)". The enforcement logic is distributed across helpers/render.py (video processing, fades, subtitles), helpers/transcribe.py (ASR configuration and caching), and helpers/pack_transcripts.py (word-boundary generation).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →