How to Interpret Frame Extraction Metadata (Engine, Candidates, Deduped Count) for Debugging in claude-video
Extract three key numbers from claude-video's metadata JSON—candidate_count, deduped_count, and selected_count—to diagnose why your frame extraction produced too few or too many frames.
When you run the watch skill in bradautomates/claude-video, the frame extraction engine prints a compact metadata object that reveals exactly what happened during processing. Understanding these numbers helps you tune parameters, identify fallback behavior, and optimize your video analysis pipeline.
What the Frame Extraction Metadata Contains
Every extraction returns a dictionary with these consistent fields:
| Key | Purpose |
|---|---|
engine |
Which strategy ran: scene, keyframe, uniform, or timestamps |
candidate_count |
Raw frames detected before any filtering |
deduped_count |
Frames removed as near-duplicates |
selected_count |
Final frames written to disk |
fallback |
True if the engine reverted to uniform sampling |
In skills/watch/scripts/frames.py, the extract_scene_or_uniform function constructs this metadata at lines 146–152:
return (
final_frames,
{
"engine": "scene",
"candidate_count": len(candidates),
"deduped_count": deduped_count,
"selected_count": len(final_frames),
"fallback": False,
},
)
Where Each Engine Produces Metadata
Scene-Aware Extraction with Uniform Fallback
The primary scene detection path in extract_scene_or_uniform populates engine: "scene" when ffmpeg's scene filter detects sufficient cuts. When scene detection fails to meet SCENE_MIN_FRAMES (8 by default), the function falls back to uniform sampling and sets engine: "uniform" with fallback: True.
Keyframe Extraction with Uniform Fallback
Inside extract_keyframes, the uniform-fallback branch at lines 663–668 creates metadata with "engine": "uniform". The successful keyframe path at lines 677–682 produces "engine": "keyframe".
Timestamp-Driven Cues
The extract_at_timestamps function returns a tuple whose second element contains "engine": "timestamps" at lines 84–88. This path also includes dropped_out_of_window when timestamps fall outside --start/--end bounds.
Decoding the Four Critical Numbers
candidate_count: Raw Detection Volume
This reveals how "rich" your video is for the chosen engine:
- Scene: Number of frame-change events from ffmpeg scene detection
- Keyframe: Count of I-frames detected
- Uniform: Equals your
--max-framesbudget (no real detection) - Timestamps: Number of timestamps you specified
A low candidate_count with fallback: True indicates the engine couldn't find enough content.
deduped_count: Static Content Indicator
Perceptual deduplication runs through _dedupe_by_deltas in dedupe_perceptual (lines 66–71). The default DEDUP_THRESHOLD of 2.0 collapses frames with similar thumbnails.
High deduped_count values suggest:
- Screen recordings with static slides
- Long stationary shots
- Video with minimal motion
To retain all frames, use --no-dedup.
selected_count: Final Output Size
This results from three processing stages:
- Deduplication (if enabled)
- Even sampling to respect budget caps via
_even_sample - Fallback triggers when minimum thresholds aren't met
When selected_count equals candidate_count - deduped_count, your video hit no budget constraints.
fallback: Engine Behavior Flag
True means the primary engine failed its minimum requirements:
- Scene detection found fewer than
SCENE_MIN_FRAMEScuts - Keyframe detection found fewer than
KEYFRAME_MINI-frames
The system automatically reverted to uniform sampling to ensure usable output.
Common Debugging Scenarios
| Symptom | Metadata Signature | Fix |
|---|---|---|
| Far fewer frames than expected | selected_count << candidate_count, fallback: True |
Increase --max-frames or use richer engine (--detail balanced) |
| Massive deduplication | deduped_count ≈ candidate_count |
Accept reduction for static content, or use --no-dedup |
| Zero candidates | candidate_count: 0 |
Check --timestamps against --start/--end window |
| Wrong engine active | engine: "uniform" despite scene/keyframe request |
Verify video has enough cuts/keyframes; fallback is working correctly |
Practical Debugging Examples
Direct frames.py execution
python -m skills.watch.scripts.frames \
/path/to/video.mp4 /tmp/frames \
--fps 1.5 --resolution 640 --max-frames 50
Output metadata:
{
"engine": "uniform",
"candidate_count": 50,
"deduped_count": 12,
"selected_count": 38,
"fallback": false
}
Interpretation: 38 distinct frames from 50 candidates after removing 12 duplicates.
High-level watch command with report
python -m skills.watch.scripts.watch \
"https://youtu.be/abc123" \
--detail balanced --max-frames 80
Report line:
- **Frames:** 62 selected from 94 candidates (scene, 7 near-duplicates dropped, full range, budget 80, cap 80)
Breakdown:
94=candidate_count7=deduped_count62=selected_count
"with uniform fallback" would indicate fallback: True.
Timestamp extraction debugging
python -m skills.watch.scripts.watch \
video.mp4 \
--timestamps "00:10,00:45,01:30" \
--detail transcript
If timestamps fall outside your range window, check for dropped_out_of_window in the raw metadata from extract_at_timestamps.
Key Source Files for Deep Debugging
| File | Purpose | Critical Lines |
|---|---|---|
skills/watch/scripts/frames.py |
Metadata construction, all engines | Scene: 146–152; Keyframe: 663–668, 677–682; Timestamps: 84–88 |
skills/watch/scripts/watch.py |
Report generation from metadata | 96–100 |
skills/watch/scripts/config.py |
Default caps and thresholds | frame_cap, SCENE_MIN_FRAMES, KEYFRAME_MIN |
To adjust behavior, modify DEDUP_THRESHOLD (2.0), SCENE_THRESHOLD, or minimum frame constants directly in frames.py or override via CLI flags.
Summary
- Read the metadata JSON that every claude-video extraction returns to understand engine behavior
- Compare
candidate_counttoselected_countto detect budget constraints or aggressive deduplication - Watch for
fallback: Truewhich signals automatic engine switching when primary methods fail - Use
--no-dedup,--max-frames, and--detailto control the three levers: duplicate removal, output volume, and extraction strategy
Frequently Asked Questions
Why does my scene extraction show engine: "uniform" with fallback: true?
The scene detector found fewer than 8 cuts (SCENE_MIN_FRAMES), so it fell back to uniform sampling. Increase your frame budget with --max-frames or accept the fallback for videos with minimal scene changes.
How do I completely disable deduplication to see all candidate frames?
Add --no-dedup to your watch command. This sets deduped_count to 0 and ensures selected_count reflects only budget capping, not perceptual similarity filtering.
What's the difference between candidate_count being low versus deduped_count being high?
Low candidate_count means the engine found little inherent content (few scene cuts, sparse keyframes, or limited timestamps). High deduped_count means plenty of candidates existed, but many were visually similar. The first needs a different engine or video; the second is expected for static content.
Where does the human-readable report get these numbers from?
The watch.py orchestrator reads the metadata tuple returned by extraction functions and injects values into markdown at lines 96–100, formatting them into the "Frames:" line you see in output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →