How video-use Edits Videos Using Natural Language Conversation
video-use turns plain-text conversations with a large language model into professional video edits by having the model read compact, phrase-level transcripts and on-demand visual thumbnails instead of processing raw footage.
The browser-use/video-use repository implements a conversation-first video editing pipeline that enables complex post-production through natural language dialogue. Rather than forcing the LLM to analyze expensive raw pixels, the system translates video content into lightweight textual representations and sparse visual confirmations, allowing the model to reason about cuts, grades, and overlays efficiently while maintaining broadcast-quality output.
The Two-Layer Representation Strategy
video-use employs a dual-layer abstraction that separates audio content from visual verification, keeping token costs low while preserving editorial precision.
Audio Transcript Layer (helpers/transcribe_batch.py)
The primary input to the LLM is a compressed transcript generated by ElevenLabs Scribe. The helpers/transcribe_batch.py script processes each source file in parallel, caching speaker IDs, word-level timestamps, and non-verbal events (such as (laugh)) as JSON. The helpers/pack_transcripts.py utility then collapses these individual JSON files into a single markdown file named takes_packed.md, typically staying under ~12 KB even for multi-minute footage.
This transcript gives the LLM exact cut points and verbal cues without requiring frame-by-frame analysis, avoiding the "45M token" anti-pattern of feeding raw video into the context window.
Visual Composite Layer (helpers/timeline_view.py)
When the LLM encounters ambiguous pauses or needs visual confirmation, it requests a targeted visual composite rather than scanning the entire video. The helpers/timeline_view.py <video> <start> <end> command renders a PNG showing a filmstrip with waveform overlays and word labels for the specific time range.
This on-demand visualization—described in the README (lines 87-88)—provides the model a quick sanity check for cut boundaries without incurring the cost of full video ingestion.
The Conversation-Driven Editing Pipeline
The workflow follows a strict seven-stage pipeline encoded in SKILL.md (lines 83-99) and visualized in the repository's README diagram (Transcribe → Pack → LLM Reasons → EDL → Render → Self-Eval).
1. Inventory and Transcription
The process begins with ffprobe scanning every source file to build a technical inventory. The system then automatically invokes helpers/transcribe_batch.py to generate per-file JSON transcripts, followed by helpers/pack_transcripts.py to create the consolidated takes_packed.md that serves as the LLM's primary view of the content.
2. Pre-scan and Context Gathering
The LLM reads the packed transcript to identify filler words, obvious slips, and structural opportunities. It then enters the conversation phase—the only point where the model interacts with the user—asking targeted questions about video type, target length, aesthetic preferences, must-keep moments, subtitle styling, and animation requirements.
3. Strategy Proposal and Confirmation
Before executing any cuts, the LLM must produce a concise 4-8 sentence strategy (e.g., "cut on silences ≥ 400ms, apply the warm_cinematic grade, add three HyperFrames overlays") and receive explicit user confirmation. This requirement is enforced as Hard Rule 11 in SKILL.md: no editing actions occur without user approval of the plan.
4. Execution via helpers/render.py
Once confirmed, the LLM orchestrates helpers/render.py, which performs several operations in sequence:
- Extracts segments using
-c copyper-segment extraction followed by concatenation (Hard Rule 2) - Applies per-segment color grades via
helpers/grade.pyusing ffmpeg filter chains - Adds 30ms audio fades to prevent pops (Hard Rule 3)
- Composites animation overlays rendered by parallel sub-agents
- Burns subtitles as the final step (Hard Rule 1)
Parallel animation sub-agents are spawned via the Agent tool (Hard Rule 10), allowing each animation slot (HyperFrames, Remotion, Manim, or PIL-based) to render concurrently rather than sequentially.
5. Self-Evaluation and Quality Assurance
After rendering, the system runs timeline_view.py on the output at every cut boundary to verify no visual jumps, audio pops, or hidden subtitle artifacts occur. The pipeline allows up to three auto-fix loops; if problems persist, the LLM surfaces them to the user for manual intervention.
6. Persistence and Continuity
The final final.mp4 is saved to <videos_dir>/edit/, while the session's reasoning and decisions are appended to project.md to maintain continuity across future editing sessions.
Hard Rules and Architectural Safeguards
The SKILL.md file (lines 22-30) defines immutable constraints that guarantee production correctness:
- Subtitles are always added last (Hard Rule 1) to prevent overlay conflicts
- Cuts never split words (Hard Rule 2) to preserve dialogue coherence
- 30ms audio fades (Hard Rule 3) prevent transition pops
- Segment extraction uses
-c copyto avoid generational quality loss before concatenation
These rules ensure that even when the LLM makes incorrect decisions, the output remains technically valid and broadcast-safe.
Practical Workflow Example
The following commands demonstrate the typical user workflow:
# Install dependencies
uv sync
brew install ffmpeg yt-dlp
# Configure API keys
cp .env.example .env
# Edit .env: ELEVENLABS_API_KEY=your_key_here
# Navigate to footage directory
cd /path/to/your/videos
claude # or codex, hermes, etc.
Within the agent console, the conversation flows naturally:
> edit these into a launch video
The LLM automatically executes the inventory stage, then prompts:
I see 3 raw takes (total 2 min 13 s).
What length do you want?
Any brand colors or subtitle style you prefer?
After gathering context, it proposes:
Strategy:
• Cut on silences ≥ 400ms, keeping all laughs.
• Apply the "warm_cinematic" grade.
• Add three HyperFrames overlays (intro, product demo, outro).
• Use UPPERCASE 2-word subtitles, margin 35.
Confirm? (yes/no)
Upon confirmation, the system executes:
python helpers/render.py edit/edl.json -o edit/preview.mp4 --preview
The final deliverable appears at edit/final.mp4 after the self-evaluation loop completes.
Summary
- video-use enables natural language video editing by translating footage into compact transcripts (
takes_packed.md) and on-demand visual thumbnails rather than raw pixels. - The two-layer representation combines word-level Scribe output from
helpers/transcribe_batch.pywith targeted PNG renders fromhelpers/timeline_view.py. - A strict conversation flow requires inventory, pre-scan, user context gathering, and confirmed strategy before
helpers/render.pyexecutes cuts. - Hard Rules enforce technical correctness: subtitles last, word-level cuts only, 30ms fades, and parallel animation sub-agents.
- The modular helper architecture (
transcribe_batch.py,pack_transcripts.py,render.py,grade.py) keeps the LLM prompt minimal while maintaining production-grade output quality.
Frequently Asked Questions
How does video-use avoid the high token costs of processing raw video?
video-use never feeds raw frames to the LLM. Instead, it uses helpers/transcribe_batch.py to create a text transcript of the audio and helpers/timeline_view.py to generate sparse PNG snapshots only when visual confirmation is needed. This keeps the context window under ~12 KB of text rather than millions of video tokens.
What are the "Hard Rules" and why do they matter?
The Hard Rules are immutable constraints defined in SKILL.md (lines 22-30) that guarantee technical validity regardless of LLM errors. Key rules include burning subtitles last (Hard Rule 1), never splitting words with cuts (Hard Rule 2), and adding 30ms audio fades (Hard Rule 3). These safeguards ensure the output is always broadcast-safe even if the model's editorial decisions are imperfect.
How does video-use handle complex animations?
The system spawns parallel animation sub-agents via the Agent tool (Hard Rule 10). Each animation slot runs independently using HyperFrames, Remotion, Manim, or PIL, so wall-time is limited by the slowest animation rather than a sequential rendering chain. The completed animations are then composited by helpers/render.py during the final output stage.
Can video-use edit videos without user confirmation?
No. According to Hard Rule 11 in SKILL.md, the LLM must propose a strategy and receive explicit user confirmation before executing any cuts. This conversational checkpoint ensures the editorial direction aligns with creator intent before irreversible rendering operations begin.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →