# Video-Use Workflow Architecture: A Deep Dive into SKILL.md

> Explore the SKILL.md video-use workflow architecture. Discover a linear eight-stage pipeline for audio-first video editing with hard rules and automated quality gates.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: architecture
- Published: 2026-06-30

---

**The [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) file in the browser-use/video-use repository defines a strictly linear, eight-stage conversation-driven pipeline for audio-first video editing that enforces production correctness through immutable "hard rules" and automated quality gates.**

The browser-use/video-use repository implements an LLM-orchestrated video editing system controlled by [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md). This master specification document outlines a reproducible workflow architecture that separates creative decision-making from technical execution while maintaining strict broadcast-quality standards. Understanding this architecture reveals how the system balances automated processing with mandatory human confirmation points.

## The Eight-Stage Pipeline

According to lines 85-99 of [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md), the workflow architecture progresses through a linear sequence of eight distinct stages with explicit feedback loops after preview and evaluation phases.

### Stage 1: Inventory

The pipeline begins by building a complete analytical view of source materials. The system runs `ffprobe` on every clip to extract technical metadata, executes [`transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/transcribe_batch.py) for parallel audio transcription, and generates [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md)—a packed transcript format that serves as the primary text representation for LLM reasoning. The [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) helper creates visual context for rapid human inspection.

### Stage 2: Pre-scan

Before invoking expensive LLM reasoning, the system scans [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) for verbal slips, mis-speaks, or unwanted phrasing. This stage records obvious issues in the editor brief, preventing simple errors from contaminating the editing strategy.

### Stage 3: Converse

The LLM describes the material in plain English and asks tailored questions regarding output type, target length, aesthetic preferences, must-keep moments, and technical requirements (animation engines, color grading, subtitle styles). This conversation establishes the creative constraints for the session.

### Stage 4: Propose Strategy

The system produces a concise, human-readable plan summarizing video shape, take selection, cut direction, animation plan, grading approach, subtitle style, and length estimate in 4-8 sentences. **Rule 11** mandates that the workflow waits for explicit user confirmation before proceeding, creating a hard boundary between planning and implementation.

### Stage 5: Execute

Once approved, the system converts the plan into concrete assets. This stage generates [`edl.json`](https://github.com/browser-use/video-use/blob/main/edl.json) via the editor sub-agent brief, runs [`render.py`](https://github.com/browser-use/video-use/blob/main/render.py) for per-segment extracts with overlays and subtitles, builds animations in parallel sub-agents (HyperFrames, Remotion, Manim, or PIL), and applies per-segment color grading via [`grade.py`](https://github.com/browser-use/video-use/blob/main/grade.py). This stage enforces **Rules 1-4 and 6-10**, which govern subtitle ordering, lossless concatenation, 30ms audio fade curves, overlay timing, cut boundaries, caching, and parallel agent coordination.

### Stage 6: Preview

The `render.py --preview` command produces a low-resolution draft for rapid visual verification. This lightweight render allows quick validation of pacing and sequencing without the computational cost of full-quality output.

### Stage 7: Self-Eval

Before presenting output to the user, the system runs automated validation. The [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) helper analyzes the rendered output around every cut (±1.5 seconds) to detect visual discontinuities, audio pops, hidden subtitles, or misaligned overlays. The evaluation limits itself to three passes; persistent failures surface to the user rather than looping indefinitely. This stage enforces **Rules 1-4 and 7-9** regarding subtitle placement, audio fades, overlay timing, and cut padding.

### Stage 8: Iterate + Persist

The system accepts natural-language feedback, re-plans, and re-renders while **never re-transcribing** (preserving the original audio analysis). Each session appends a structured summary to [`project.md`](https://github.com/browser-use/video-use/blob/main/project.md), creating an audit trail of creative decisions and technical parameters.

## Hard Rules and Production Constraints

Lines 22-30 of [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) define immutable "hard rules" that guarantee technical correctness regardless of creative decisions:

- **Subtitle ordering**: Subtitles render last in the composite stack to prevent occlusion by overlays
- **Lossless concatenation**: Segments concatenate without generational loss between cuts
- **Audio fade curves**: All cuts apply 30ms audio fade envelopes to prevent pops
- **Overlay timing**: Animation overlays synchronize to frame-accurate boundaries
- **Cut padding**: Self-evaluation inspects ±1.5 seconds around each cut boundary

These rules apply universally during the Execute and Self-Eval stages, ensuring that LLM-generated creative decisions never violate technical broadcast standards.

## Directory Structure and Helper Scripts

The workflow expects a specific directory layout defined in lines 41-55 of [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md). All artifacts live under `<videos_dir>/edit/`, with helper scripts invoked relative to the [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) location.

Key implementation files referenced in lines 72-80 include:

- **[`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py)**: Single-file Scribe transcription with caching
- **[`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py)**: Parallel batch processing for multiple clips
- **[`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py)**: Converts raw JSON transcripts into phrase-level [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md)
- **[`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py)**: Core rendering engine handling extraction, concatenation, overlay compositing, and subtitle burn-in
- **[`helpers/grade.py`](https://github.com/browser-use/video-use/blob/main/helpers/grade.py)**: Applies color-grade filters per segment
- **[`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py)**: Generates film-strip and waveform PNGs for visual inspection

## Practical Implementation: Running the Workflow

Below are command-line implementations for each architectural stage. All paths are relative to the directory containing [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md).

### Stage 1: Inventory Generation

```bash

# Inspect source metadata

ffprobe *.mp4

# Parallel transcription

transcribe_batch.py .

# Pack transcripts for LLM consumption

pack_transcripts.py --edit-dir edit

# Generate visual timeline

timeline_view.py clip1.mp4 0 10

```

### Stage 5: Execution Pipeline

```bash

# Render with subtitles and overlays

render.py edit/edl.json -o edit/render.mp4 --build-subtitles

# Apply cinematic color grade per segment

grade.py edit/clips_graded/segment1.mp4 \
  -o edit/graded/segment1.mp4 \
  --preset warm_cinematic

# Create animation slot via HyperFrames

mkdir -p edit/animations/slot_1
cd edit/animations/slot_1
npx --yes hyperframes init . --example blank --non-interactive --skip-skills
npx --yes hyperframes render . -o render.mp4

```

### Stage 6-7: Preview and Validation

```bash

# Generate low-res preview

render.py edit/edl.json -o edit/preview.mp4 --preview

# Validate cuts (example: inspecting cut at 12.3s)

timeline_view.py edit/preview.mp4 11.8 12.8

```

### Stage 8: Session Persistence

```bash

# Append session log to project history

cat >> edit/project.md <<'EOF'

## Session 3 — 2026-06-30

**Strategy:** 3-minute tech launch, warm grade, HyperFrames animation at 0:45

**Decisions:** Cut mis-speak at 0:23, extend CTA by 2s

**Reasoning log:** User confirmed strategy via Rule 11 checkpoint
EOF

```

## Summary

- **[`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) defines an eight-stage linear pipeline** with feedback loops after preview and self-evaluation stages.
- **Rule 11 mandates strategy confirmation** before execution, creating an immutable boundary between planning and rendering.
- **Hard rules 1-10 enforce technical correctness** regarding subtitle ordering, lossless concatenation, audio fades, and overlay timing.
- **Self-evaluation runs automated quality gates** using `timeline_view` to detect visual and audio artifacts before user review.
- **The architecture never re-transcribes** during iteration, preserving the original audio analysis while allowing unlimited re-renders.

## Frequently Asked Questions

### What makes the Video-Use workflow "audio-first"?

The workflow prioritizes spoken content as the primary editing reference. The [`transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/transcribe_batch.py) and [`pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/pack_transcripts.py) helpers convert all audio into structured [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) transcripts before any visual editing occurs. Cuts align to phrase boundaries and verbal cues rather than purely visual markers, ensuring the narrative flow drives the editing decisions.

### Can I skip stages in the SKILL.md workflow?

No. The architecture enforces strict linear progression through all eight stages. While feedback loops allow returning to earlier stages after Preview or Self-Eval, the system cannot bypass Inventory (which builds the transcript corpus) or Propose Strategy (which requires user confirmation under Rule 11). Skipping Pre-scan would allow obvious verbal errors to contaminate the editing strategy.

### How does the self-evaluation stage detect editing errors?

The Self-Eval stage runs [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) on the rendered output, inspecting ±1.5 seconds around every cut boundary defined in [`edl.json`](https://github.com/browser-use/video-use/blob/main/edl.json). It checks for visual discontinuities, audio waveform pops that violate the 30ms fade rule, subtitles hidden by overlay layers, and misaligned animation timing. The system attempts automatic correction for up to three passes before surfacing persistent issues to the user.

### What happens if I need to change the video after rendering?

The workflow supports natural-language iteration without restarting. If you request changes after Preview or Self-Eval, the system re-enters the Converse stage, generates a new strategy (subject to Rule 11 confirmation), and re-executes the render pipeline. Importantly, the system never re-runs [`transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/transcribe_batch.py), preserving the original audio analysis in [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) while updating the [`edl.json`](https://github.com/browser-use/video-use/blob/main/edl.json) and rendered outputs.