What Do Claude-Video’s `--detail` Modes Do? (efficient, balanced, token-burner, transcript)
The --detail flag controls frame extraction density, offering four modes: efficient (minimal keyframes), balanced (default scene-aware sampling), token-burner (up to 250 frames), and transcript (audio-only, no frames).
The bradautomates/claude-video repository provides a watch skill for processing video content. These --detail modes determine exactly how many visual frames are sampled from the source material, directly impacting both token consumption and the granularity of AI analysis.
How --detail Modes Map to Frame Extraction
In skills/watch/scripts/config.py (lines 51-73), the DETAILS list and frame_cap() function define the behavior for each mode. The selection logic is then executed in skills/watch/scripts/watch.py (lines 66-72 and 72-333), where the code checks if detail == "efficient", if detail == "balanced", and so on to assign the appropriate engine and budget.
efficient: Keyframe-Only Extraction
The efficient mode employs the keyframes engine (engine_label = "keyframes"). It extracts only the most essential frames—typically those representing major scene changes—keeping the frame budget minimal. This mode is ideal when you need a quick, token-efficient summary with just a handful of representative stills.
balanced: Scene-Aware Sampling (Default)
The balanced mode is the default, as defined by DEFAULT_DETAIL = "balanced" in skills/watch/scripts/config.py. It uses the scene-aware frames engine to sample more densely than efficient while still respecting the user-specified frame cap. This provides the best trade-off between visual coverage and token cost for most workflows.
token-burner: Maximum Frame Retention
The token-burner mode disables the standard frame budget cap, instructing the engine to retain all detected scene-change frames. However, skills/watch/scripts/watch.py enforces a hard warning limit of 250 frames (if len(frames) > 250). Use this mode when you need maximum visual detail and are willing to consume significantly more tokens.
transcript: Audio-Only Processing
The transcript mode skips frame extraction entirely, returning only the Whisper-generated transcript or captions. If no transcript is available, the system automatically falls back to the balanced mode. This is optimal for audio-centric analysis or when visual content is irrelevant.
Configuration and Environment Overrides
You can inspect the default configuration in skills/watch/scripts/config.py, where DEFAULT_DETAIL is explicitly set to "balanced". The frame_cap() function maps each string value to its corresponding frame limit.
To override the default without typing the flag on every command, set the WATCH_DETAIL environment variable:
export WATCH_DETAIL=efficient
This variable is read at runtime by the configuration loader, allowing you to change the default behavior system-wide.
CLI Usage Examples
Run the watch skill with explicit detail flags to control extraction density:
# Minimal frames, lowest token cost
watch https://example.com/video.mp4 --detail efficient
# Default behavior (explicit)
watch https://example.com/video.mp4 --detail balanced
# Maximum frames (up to 250 limit)
watch https://example.com/video.mp4 --detail token-burner
# Audio only, no frames extracted
watch https://example.com/video.mp4 --detail transcript
Summary
efficient: Uses thekeyframesengine for minimal scene-change frames only.balanced: Defaultscene-aware framesengine offering moderate density; controlled byDEFAULT_DETAILinskills/watch/scripts/config.py.token-burner: Removes frame caps but enforces a hard limit of 250 frames inskills/watch/scripts/watch.py.transcript: Disables frame extraction, returning only Whisper transcripts with automatic fallback tobalancedif audio is unavailable.- Configuration: Modify defaults via the
WATCH_DETAILenvironment variable or editskills/watch/scripts/config.py.
Frequently Asked Questions
What is the default --detail mode if I don't specify one?
According to tests/test_config.py and the source in skills/watch/scripts/config.py, the default mode is balanced. This is enforced by the DEFAULT_DETAIL constant and validated in the configuration test suite.
How do I permanently set a different default detail mode?
Set the WATCH_DETAIL environment variable to your preferred mode (e.g., efficient, token-burner). The configuration loader in skills/watch/scripts/config.py reads this variable at runtime, allowing you to override the built-in default without modifying CLI commands.
What happens when I use --detail transcript on a video without audio?
If no Whisper-generated transcript is available, the code in skills/watch/scripts/watch.py automatically falls back to balanced mode and extracts scene-aware frames instead. You will receive visual analysis rather than an empty response.
Is there a safety limit on frames in token-burner mode?
Yes. While token-burner disables the standard frame budget cap, skills/watch/scripts/watch.py enforces a hard warning limit of 250 frames. If the video contains more scene changes than this threshold, the system triggers a warning but continues processing up to that cap.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →