# How video-use Edits Videos Using Natural Language Conversation

> Discover how video-use edits videos using natural language conversations. Learn how it transforms text prompts into professional video edits without processing raw footage.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: how-to-guide
- Published: 2026-07-01

---

**video-use turns plain-text conversations with a large language model into professional video edits by having the model read compact, phrase-level transcripts and on-demand visual thumbnails instead of processing raw footage.**

The `browser-use/video-use` repository implements a **conversation-first video editing pipeline** that enables complex post-production through natural language dialogue. Rather than forcing the LLM to analyze expensive raw pixels, the system translates video content into lightweight textual representations and sparse visual confirmations, allowing the model to reason about cuts, grades, and overlays efficiently while maintaining broadcast-quality output.

## The Two-Layer Representation Strategy

video-use employs a dual-layer abstraction that separates audio content from visual verification, keeping token costs low while preserving editorial precision.

### Audio Transcript Layer ([`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py))

The primary input to the LLM is a compressed transcript generated by ElevenLabs Scribe. The [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) script processes each source file in parallel, caching speaker IDs, word-level timestamps, and non-verbal events (such as `(laugh)`) as JSON. The [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) utility then collapses these individual JSON files into a single markdown file named [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md), typically staying under ~12 KB even for multi-minute footage.

This transcript gives the LLM **exact cut points** and verbal cues without requiring frame-by-frame analysis, avoiding the "45M token" anti-pattern of feeding raw video into the context window.

### Visual Composite Layer ([`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py))

When the LLM encounters ambiguous pauses or needs visual confirmation, it requests a targeted visual composite rather than scanning the entire video. The `helpers/timeline_view.py <video> <start> <end>` command renders a PNG showing a filmstrip with waveform overlays and word labels for the specific time range.

This on-demand visualization—described in the README (lines 87-88)—provides the model a quick sanity check for cut boundaries without incurring the cost of full video ingestion.

## The Conversation-Driven Editing Pipeline

The workflow follows a strict seven-stage pipeline encoded in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) (lines 83-99) and visualized in the repository's README diagram (`Transcribe → Pack → LLM Reasons → EDL → Render → Self-Eval`).

### 1. Inventory and Transcription

The process begins with `ffprobe` scanning every source file to build a technical inventory. The system then automatically invokes [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) to generate per-file JSON transcripts, followed by [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) to create the consolidated [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) that serves as the LLM's primary view of the content.

### 2. Pre-scan and Context Gathering

The LLM reads the packed transcript to identify filler words, obvious slips, and structural opportunities. It then enters the **conversation phase**—the only point where the model interacts with the user—asking targeted questions about video type, target length, aesthetic preferences, must-keep moments, subtitle styling, and animation requirements.

### 3. Strategy Proposal and Confirmation

Before executing any cuts, the LLM must produce a concise 4-8 sentence strategy (e.g., "cut on silences ≥ 400ms, apply the *warm_cinematic* grade, add three HyperFrames overlays") and receive explicit user confirmation. This requirement is enforced as **Hard Rule 11** in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md): no editing actions occur without user approval of the plan.

### 4. Execution via [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py)

Once confirmed, the LLM orchestrates [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), which performs several operations in sequence:
- Extracts segments using `-c copy` per-segment extraction followed by concatenation (Hard Rule 2)
- Applies per-segment color grades via [`helpers/grade.py`](https://github.com/browser-use/video-use/blob/main/helpers/grade.py) using ffmpeg filter chains
- Adds 30ms audio fades to prevent pops (Hard Rule 3)
- Composites animation overlays rendered by parallel sub-agents
- Burns subtitles as the final step (Hard Rule 1)

Parallel animation sub-agents are spawned via the `Agent` tool (Hard Rule 10), allowing each animation slot (HyperFrames, Remotion, Manim, or PIL-based) to render concurrently rather than sequentially.

### 5. Self-Evaluation and Quality Assurance

After rendering, the system runs [`timeline_view.py`](https://github.com/browser-use/video-use/blob/main/timeline_view.py) on the output at every cut boundary to verify no visual jumps, audio pops, or hidden subtitle artifacts occur. The pipeline allows up to three auto-fix loops; if problems persist, the LLM surfaces them to the user for manual intervention.

### 6. Persistence and Continuity

The final `final.mp4` is saved to `<videos_dir>/edit/`, while the session's reasoning and decisions are appended to [`project.md`](https://github.com/browser-use/video-use/blob/main/project.md) to maintain continuity across future editing sessions.

## Hard Rules and Architectural Safeguards

The [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) file (lines 22-30) defines immutable constraints that guarantee production correctness:
- **Subtitles are always added last** (Hard Rule 1) to prevent overlay conflicts
- **Cuts never split words** (Hard Rule 2) to preserve dialogue coherence
- **30ms audio fades** (Hard Rule 3) prevent transition pops
- **Segment extraction uses `-c copy`** to avoid generational quality loss before concatenation

These rules ensure that even when the LLM makes incorrect decisions, the output remains technically valid and broadcast-safe.

## Practical Workflow Example

The following commands demonstrate the typical user workflow:

```bash

# Install dependencies

uv sync
brew install ffmpeg yt-dlp

# Configure API keys

cp .env.example .env

# Edit .env: ELEVENLABS_API_KEY=your_key_here

# Navigate to footage directory

cd /path/to/your/videos
claude  # or codex, hermes, etc.

```

Within the agent console, the conversation flows naturally:

```text
> edit these into a launch video

```

The LLM automatically executes the inventory stage, then prompts:

```text
I see 3 raw takes (total 2 min 13 s).
What length do you want?
Any brand colors or subtitle style you prefer?

```

After gathering context, it proposes:

```text
Strategy:
• Cut on silences ≥ 400ms, keeping all laughs.
• Apply the "warm_cinematic" grade.
• Add three HyperFrames overlays (intro, product demo, outro).
• Use UPPERCASE 2-word subtitles, margin 35.
Confirm? (yes/no)

```

Upon confirmation, the system executes:

```bash
python helpers/render.py edit/edl.json -o edit/preview.mp4 --preview

```

The final deliverable appears at `edit/final.mp4` after the self-evaluation loop completes.

## Summary

- **video-use** enables natural language video editing by translating footage into compact transcripts ([`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md)) and on-demand visual thumbnails rather than raw pixels.
- The **two-layer representation** combines word-level Scribe output from [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) with targeted PNG renders from [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py).
- A **strict conversation flow** requires inventory, pre-scan, user context gathering, and confirmed strategy before [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) executes cuts.
- **Hard Rules** enforce technical correctness: subtitles last, word-level cuts only, 30ms fades, and parallel animation sub-agents.
- The modular helper architecture ([`transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/transcribe_batch.py), [`pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/pack_transcripts.py), [`render.py`](https://github.com/browser-use/video-use/blob/main/render.py), [`grade.py`](https://github.com/browser-use/video-use/blob/main/grade.py)) keeps the LLM prompt minimal while maintaining production-grade output quality.

## Frequently Asked Questions

### How does video-use avoid the high token costs of processing raw video?

video-use never feeds raw frames to the LLM. Instead, it uses [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) to create a text transcript of the audio and [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) to generate sparse PNG snapshots only when visual confirmation is needed. This keeps the context window under ~12 KB of text rather than millions of video tokens.

### What are the "Hard Rules" and why do they matter?

The Hard Rules are immutable constraints defined in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) (lines 22-30) that guarantee technical validity regardless of LLM errors. Key rules include burning subtitles last (Hard Rule 1), never splitting words with cuts (Hard Rule 2), and adding 30ms audio fades (Hard Rule 3). These safeguards ensure the output is always broadcast-safe even if the model's editorial decisions are imperfect.

### How does video-use handle complex animations?

The system spawns parallel animation sub-agents via the `Agent` tool (Hard Rule 10). Each animation slot runs independently using HyperFrames, Remotion, Manim, or PIL, so wall-time is limited by the slowest animation rather than a sequential rendering chain. The completed animations are then composited by [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) during the final output stage.

### Can video-use edit videos without user confirmation?

No. According to Hard Rule 11 in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md), the LLM must propose a strategy and receive explicit user confirmation before executing any cuts. This conversational checkpoint ensures the editorial direction aligns with creator intent before irreversible rendering operations begin.