Claude-Video Detail Modes Explained: efficient vs balanced vs token-burner vs transcript
The four --detail modes in Claude-Video control how many frames are extracted from a video: efficient uses keyframes only, balanced (default) uses scene-aware sampling, token-burner keeps all scene-change frames up to 250, and transcript skips frames entirely to return only the Whisper transcript.
Claude-Video's watch skill offers granular control over visual frame extraction through the --detail flag. Each mode determines the underlying frame-selection engine, the resulting token budget, and whether video analysis prioritizes speed, coverage, or audio-only content. Understanding these differences ensures you choose the right trade-off between information density and API cost.
How Detail Modes Map to Frame Engines
The DETAILS configuration in skills/watch/scripts/config.py defines four string constants that drive distinct behavior in skills/watch/scripts/watch.py. Here's how each mode operates at the code level:
efficient Mode: Keyframe-Only Extraction
- Engine label:
"keyframes" - Behavior: Selects only the most essential frames—typically major scene-change moments.
- Frame budget: Lowest among visual modes; minimizes token consumption by aggressively filtering duplicates.
Use efficient when you need rapid, low-cost summaries and can tolerate sparse visual coverage.
watch https://example.com/video.mp4 --detail efficient
balanced Mode: Scene-Aware Sampling (Default)
- Engine label:
"scene-aware frames" - Behavior: Samples frames more densely than
efficientwhile respecting the user-specified frame cap. - Frame budget: Moderate; the default
DEFAULT_DETAILvalue.
This mode is defined in skills/watch/scripts/config.py as the fallback when no --detail argument is provided, and confirmed by tests/test_config.py lines 7-13.
watch https://example.com/video.mp4 # uses balanced by default
watch https://example.com/video.mp4 --detail balanced
token-burner Mode: Maximum Visual Coverage
- Engine label: Scene-aware engine with disabled or dramatically raised frame cap
- Behavior: Retains all scene-change frames, subject to a hard warning limit of 250 frames:
if len(frames) > 250 - Frame budget: Highest; deliberately trades token efficiency for comprehensive visual context.
Select token-burner when the analysis demands maximum visual fidelity and you accept higher API costs.
watch https://example.com/video.mp4 --detail token-burner
transcript Mode: Audio-Only Analysis
- Engine label: None (frame extraction skipped entirely)
- Behavior: Returns only the Whisper-generated transcript or captions. If no transcript is available, the code falls back to
balancedframes. - Frame budget: Zero (unless fallback occurs).
Ideal for audio-centric workflows or when visual content is irrelevant.
watch https://example.com/video.mp4 --detail transcript
Configuration and Environment Overrides
The frame_cap(detail) function in skills/watch/scripts/config.py (lines 51-73) maps each mode to its numeric frame budget. You can override the default mode without passing --detail on every command:
export WATCH_DETAIL=efficient
watch https://example.com/video.mp4 # runs in efficient mode
Environment precedence is validated in tests/test_config.py.
Core Implementation Files
| File | Purpose |
|---|---|
skills/watch/scripts/config.py |
Defines DETAILS list, DEFAULT_DETAIL = "balanced", and frame_cap(detail) budget mapping |
skills/watch/scripts/watch.py |
Parses --detail argument, selects engine label, applies caps, and issues 250-frame warnings |
tests/test_config.py |
Unit tests for default mode and environment variable overrides |
skills/watch/scripts/setup.py |
Reports watch_detail: "balanced" as the configured default |
Summary
efficient→ keyframe engine, minimal tokens, fastest executionbalanced→ scene-aware engine, moderate coverage, default behaviortoken-burner→ all scene-change frames, up to 250, maximum token usagetranscript→ Whisper audio only, withbalancedfallback when unavailable
Each mode is implemented in skills/watch/scripts/config.py and executed through skills/watch/scripts/watch.py, giving you precise control over Claude-Video's visual analysis pipeline.
Frequently Asked Questions
What happens if I don't specify a --detail mode?
Claude-Video defaults to balanced mode. This is hard-coded as DEFAULT_DETAIL = "balanced" in skills/watch/scripts/config.py and verified by test assertions in tests/test_config.py.
Will token-burner mode ever exceed 250 frames?
No. The watch.py implementation enforces a hard warning threshold: if len(frames) > 250. While the mode disables the standard frame cap, the 250-frame safety limit prevents runaway token consumption.
Can I use --detail transcript if the video has no audio track?
Yes, but the behavior changes. When no Whisper transcript is available, the code automatically falls back to balanced frame extraction. You will receive visual frames instead of an empty response.
How do I permanently change my default detail mode?
Set the WATCH_DETAIL environment variable to any valid mode (efficient, balanced, token-burner, or transcript). This overrides DEFAULT_DETAIL without modifying the source code or passing --detail on every command.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →