Near-Duplicate Frame Removal Algorithm in Claude-Video: Implementation and Usage
Claude-Video removes visually similar frames by comparing downscaled thumbnails and dropping frames whose RGB distance to the last kept frame falls below a configurable threshold, while the --no-dedup flag bypasses this step entirely to preserve every extracted frame.
Claude-Video is an open-source video analysis tool that extracts frames from video sources and optimizes processing by collapsing near-identical images. The near-duplicate frame removal algorithm, implemented in skills/watch/scripts/frames.py, uses perceptual hashing techniques to reduce redundant data before frame sampling. Understanding this deduplication logic helps developers tune performance and debug frame selection issues.
How the Near-Duplicate Frame Removal Algorithm Works
The algorithm operates in four distinct phases within the dedupe_perceptual function in skills/watch/scripts/frames.py:
Downscaling to Thumbnails
Every extracted frame is resized to a tiny thumbnail of size DEDUP_THUMB × DEDUP_THUMB pixels. This constant is defined near the top of skills/watch/scripts/frames.py and reduces the computational cost of per-pixel comparisons. By working with scaled-down versions rather than full-resolution images, the system can process large video files efficiently.
Computing Frame-to-Frame Deltas
The algorithm calculates the Euclidean (L2) distance between the RGB values of consecutive thumbnails. This distance metric quantifies the visual difference between adjacent frames in the video sequence.
Greedy Drop Logic
The core deduplication routine _dedupe_by_deltas (lines 479-527) implements a greedy filtering approach:
- It keeps the first frame automatically
- For each subsequent frame, it calculates the delta to the last kept thumbnail
- If the delta exceeds the configurable
threshold(default 2.0), the frame is kept - Otherwise, the frame is dropped as a near-duplicate
This ensures that retained frames represent distinct visual states while removing shots with minimal motion or static scenes.
Metadata Integration
After processing, the count of dropped frames is stored in the metadata under the key deduped_count (lines 747-753). The dedup step runs automatically before the frame-budget cap is applied, meaning downstream sampling operates on the reduced set (lines 543-549).
How the --no-dedup Flag Works
The --no-dedup CLI flag disables the perceptual deduplication pipeline entirely. When invoked:
- The
watchcommand inskills/watch/scripts/watch.py(lines 212-224) parses the flag and setsdedup=False - The frame-processing pipeline skips the call to
dedupe_perceptual - Every extracted frame is preserved unchanged
- The resulting metadata shows
deduped_countequal to 0
This flag is useful when debugging frame extraction issues, processing content where exact timing matters, or analyzing videos with subtle frame-level differences that might fall below the default threshold.
Practical Code Examples
To process a video with the default near-duplicate frame removal behavior:
claude-video watch https://example.com/video.mp4
To preserve all frames and skip deduplication:
claude-video watch https://example.com/video.mp4 --no-dedup
For direct Python integration, import the frames module and control deduplication manually:
from skills.watch.scripts import frames
# candidates is a list of Path objects pointing to extracted JPEG frames
deduped_frames, dropped = frames.dedupe_perceptual(candidates)
# deduped_frames contains only kept frames, dropped indicates removal count
# To skip deduplication (equivalent to --no-dedup):
deduped_frames, dropped = (candidates, 0)
Summary
- The near-duplicate frame removal algorithm in Claude-Video uses thumbnail comparison and Euclidean distance to filter redundant frames before sampling.
- Implementation resides in
skills/watch/scripts/frames.py, specifically the_dedupe_by_deltasfunction (lines 479-527). - The default threshold of 2.0 balances sensitivity and compression, configurable via the
thresholdparameter. - Metadata tracking stores the drop count in
deduped_countfor transparency and debugging (lines 747-753). - The
--no-dedupflag inskills/watch/scripts/watch.py(lines 212-224) completely bypasses the deduplication step when visual fidelity to the original frame rate is required.
Frequently Asked Questions
What is the default threshold for considering frames as duplicates?
The default threshold is 2.0, representing the Euclidean distance between RGB values of downscaled thumbnails. You can adjust this parameter when calling dedupe_perceptual directly to make the algorithm more or less aggressive.
Where does the deduplication occur in the processing pipeline?
Deduplication runs immediately after frame extraction and before the frame-budget cap is applied (lines 543-549 in frames.py). This ensures that downstream sampling works on the reduced set, maximizing the diversity of frames within budget constraints.
Does using --no-dedup affect the final metadata output?
Yes. When --no-dedup is passed, the deduped_count metadata field will always be 0, indicating no frames were removed during the deduplication phase. The frame count will reflect every extracted frame from the video source.
What thumbnail size does Claude-Video use for comparison?
The algorithm resizes frames to DEDUP_THUMB × DEDUP_THUMB pixels, a constant defined in skills/watch/scripts/frames.py. This aggressive downscaling reduces memory usage and computational overhead while preserving enough visual information to detect meaningful changes between frames.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →