Scene-Change Detection vs. Keyframe Extraction in Efficient vs. Balanced Modes
In the claude-video repository, --detail efficient decodes only I-frames (keyframes) using fast decoder skipping, while --detail balanced performs full scene-change detection with perceptual analysis, falling back to uniform sampling for static videos.
The claude-video tool provides intelligent frame extraction through the watch skill, which offers multiple detail levels for analyzing video content. Understanding how the --detail efficient and --detail balanced modes differ is essential for optimizing between processing speed and visual coverage when preparing videos for AI analysis.
How Efficient Mode Works (Keyframe Extraction)
The efficient detail mode prioritizes speed by leveraging the video encoder's existing keyframe structure.
The Keyframe Engine Implementation
In skills/watch/scripts/frames.py, the extract_keyframes function implements the efficient engine by invoking ffmpeg with -skip_frame nokey【link:/skills/watch/scripts/frames.py#L776-L842】:
cmd = [
"ffmpeg",
"-skip_frame", "nokey",
"-i", video_path,
"-vf", f"{_scale_filter(resolution)},showinfo",
"-vsync", "vfr",
"-q:v", "4",
output_pattern,
]
This command instructs the decoder to reconstruct only I-frames (keyframes), which encoders typically insert at scene cuts. The approach is computationally cheap but provides a coarse representation that may miss subtle transitions between keyframes.
Processing Pipeline and Fallbacks
After extraction, the pipeline applies two optimization steps:
- De-duplication: Near-identical frames are removed using
dedupe_perceptual - Even sampling: The
_even_samplefunction ensures the final count respects the 50-frame cap
If fewer than KEYFRAME_MIN (4) keyframes are found, the engine falls back to the uniform extractor (extract) to guarantee minimum coverage【link:/skills/watch/scripts/frames.py#L866-L878】.
How Balanced Mode Works (Scene-Change Detection)
The balanced detail mode provides higher-quality analysis by performing full scene-change detection across the entire video timeline.
Scene-Change Detection Algorithm
Implemented in skills/watch/scripts/frames.py within extract_scene_or_uniform【link:/skills/watch/scripts/frames.py#L1010-L1054】, this mode uses extract_scene_candidates to run ffmpeg with a scene detection filter:
# Scene detection threshold
SCENE_THRESHOLD = 0.20
The filter analyzes the video during a full decode pass and emits a frame each time the scene-change score exceeds 0.20【link:/skills/watch/scripts/frames.py#L251-L258】. This captures every perceptual cut regardless of encoder keyframe placement.
Fallback to Uniform Sampling
If the detected scene count falls below SCENE_MIN_FRAMES (8), the video is considered static and the engine switches to the uniform sampler (extract), which extracts frames at fixed intervals determined by auto_fps calculations. After scene detection, the pipeline applies the same de-duplication and even-sampling logic to respect the 100-frame cap.
Key Differences Compared
| Aspect | Efficient (Keyframes) | Balanced (Scene-Change) |
|---|---|---|
| Performance | Fast – skips most frames with single-pass decoding | Slower – requires full decode plus scene analysis |
| Granularity | Coarse – only encoder-chosen cuts | Fine – detects any change exceeding perceptual threshold |
| Frame Cap | 50 frames (default) | 100 frames (default) |
| Fallback Trigger | < 4 keyframes found | < 8 scene cuts detected |
| Engine Function | extract_keyframes |
extract_scene_or_uniform |
Practical Usage Examples
Run in efficient mode for quick previews:
watch https://youtu.be/xyz123 --detail efficient --resolution 720
The resulting metadata shows "engine": "keyframe" with "selected_count" ≤ 50.
Run in balanced mode for detailed analysis:
watch https://youtu.be/xyz123 --detail balanced --resolution 720
The output shows "engine": "scene" (or "uniform" if static) with up to 100 frames.
Inspect engine selection programmatically:
from frames import extract_keyframes, extract_scene_or_uniform
# Efficient mode
frames, meta = extract_keyframes("clip.mp4", Path("/tmp/out"), max_frames=50)
print(meta["engine"]) # → keyframe
# Balanced mode
frames, meta = extract_scene_or_uniform(
"clip.mp4", Path("/tmp/out"),
fps=1.0, target_frames=100,
max_frames=100, dedup=False,
)
print(meta["engine"]) # → scene (or uniform)
The selection logic in skills/watch/scripts/watch.py (lines 504-527) dispatches to the appropriate engine based on the --detail flag【link:/skills/watch/scripts/watch.py#L504-L527】.
Summary
- Efficient mode uses
extract_keyframesto decode only I-frames with-skip_frame nokey, capping at 50 frames and falling back to uniform sampling if fewer than 4 keyframes exist. - Balanced mode uses
extract_scene_or_uniformto perform full scene-change detection with a threshold of0.20, capping at 100 frames and falling back to uniform sampling when fewer than 8 scene cuts are detected. - Both modes apply
dedupe_perceptualand_even_sampleto optimize the final frame selection within their respective budgets.
Frequently Asked Questions
When should I use efficient mode versus balanced mode?
Use --detail efficient when processing speed and token economy are priorities, such as for quick video previews or long content where coarse sampling is acceptable. Use --detail balanced when you need comprehensive visual coverage for detailed analysis, as it captures every perceptual scene change up to the 100-frame limit.
Why does balanced mode fall back to uniform sampling?
According to the source code in skills/watch/scripts/frames.py, balanced mode falls back to uniform sampling when fewer than SCENE_MIN_FRAMES (8) scene cuts are detected. This prevents sparse coverage in static videos like lectures or slideshows where scene-change detection would return insufficient frames.
What happens if a video has no scene cuts in efficient mode?
If extract_keyframes finds fewer than KEYFRAME_MIN (4) keyframes, it automatically falls back to the uniform extractor (extract) to ensure minimum coverage. This guarantees that even compressed videos with few I-frames still provide analyzable content.
How do the frame caps affect extraction quality?
The 50-frame cap in efficient mode and 100-frame cap in balanced mode are enforced by _even_sample after de-duplication. These limits prevent token overflow while the even-sampling algorithm distributes frames chronologically across the video duration, ensuring temporal coverage rather than clustering at the beginning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →