Why Do Long Videos Show Sparse Frame Coverage in Claude Video: Causes and Fixes
Long videos show sparse frame coverage because the tool budgets a fixed number of frames across the entire duration, forcing the auto-fps logic to drop to the minimum 2.0 fps floor. Fix this by using --detail token-burner to remove the cap, increasing --max-frames, or zooming into a sub-range with --start and --end.
In the bradautomates/claude-video repository, the video processing pipeline automatically throttles frame extraction to manage API costs and token usage. This causes sparse frame coverage where lengthy content is represented by widely spaced screenshots. Understanding the auto_fps() logic in skills/watch/scripts/frames.py reveals exactly why this happens and how to override it.
How the Frame Budget Creates Sparse Coverage
The Fixed Frame Cap (max_frames)
The extraction logic budgets a fixed number of frames based on the selected detail level. In "balanced" mode, max_frames defaults to 100. The system calculates the target frame count using the formula found at lines 22-38 in skills/watch/scripts/frames.py:
target = min(max_frames, round(fps * duration))
This equation ensures the total extracted frames never exceeds the budget, regardless of video length.
Auto-FPS Floor Limitations
The auto_fps() function divides the max_frames budget by the video duration to determine sampling density. For videos exceeding roughly 50 seconds, this calculation drops below the MAX_FPS = 2.0 constant defined in frames.py. When this occurs, the system forces extraction at 2 frames per second, creating large gaps between captured moments in long videos.
Scene Detection Fallbacks
When scene detection yields fewer than SCENE_MIN_FRAMES = 8 cuts, the logic falls back to uniform sampling (lines 49-57 in frames.py). If the frame budget is already exhausted due to length, this fallback results in even sparser coverage as the system distributes the remaining frames across the full duration.
Warning Triggers
For videos longer than 10 minutes processed with default caps, watch.py generates explicit warnings (lines 26-34) indicating that frame coverage may be insufficient for accurate analysis.
Six Methods to Fix Sparse Frame Coverage
1. Use Uncapped Token-Burner Mode
--detail token-burner disables the max_frames budget entirely. This mode extracts every detected scene change without the 100-frame constraint, providing complete visual coverage regardless of video length.
2. Increase the Frame Cap
Raise --max-frames N to provide more samples across the same duration. Increasing from the default 100 to 300 triples your frame density, though API costs scale accordingly.
3. Zoom Into Specific Ranges
Use --start and --end timestamps to reduce effective duration. This concentrates the fixed budget on a shorter span, effectively increasing local fps without raising the global cap.
4. Override Auto-FPS
Force a specific extraction rate with --fps N. Note that this remains constrained by the MAX_FPS = 2.0 ceiling in the source code unless you manually edit that constant in frames.py.
5. Select Alternative Detail Levels
--detail efficient: Uses keyframes only, often resulting in fewer frames than "balanced" mode.--detail balanced: Default scene-detection mode with the 100-frame cap and warnings for long videos.--detail token-burner: Uncapped scene-detection mode with nomax_framesrestriction.
6. Configure Global Defaults
Set WATCH_DETAIL=token-burner in ~/.config/watch/.env to permanently disable caps. The config.py module reads this environment variable during initialization, applying your preference to all future invocations.
Practical Code Examples
# Extract every scene-change frame without budget limits
watch https://example.com/long-video.mp4 --detail token-burner
# Triple the frame budget for denser coverage
watch https://example.com/long-video.mp4 --max-frames 300
# Focus the budget on a specific 2-minute segment
watch https://example.com/long-video.mp4 --start 00:10:00 --end 00:12:00
# Force maximum allowed fps extraction
watch https://example.com/long-video.mp4 --fps 2.0
Implementation Details in the Source Code
The sparse coverage behavior stems from three critical components in the claude-video codebase:
skills/watch/scripts/frames.py contains the core auto_fps() logic (lines 22-38) that implements the budget calculation. It also defines SCENE_MIN_FRAMES = 8 (lines 21-30) as the threshold for falling back to uniform sampling when scene detection is insufficient.
skills/watch/scripts/watch.py (lines 26-34) implements the warning system for videos exceeding 10 minutes and coordinates the frame budget allocation between detail modes.
skills/watch/scripts/config.py defines the default max_frames values for each detail level and reads the WATCH_DETAIL environment variable to determine processing behavior without command-line flags.
Summary
- Sparse coverage occurs when
auto_fps()hits theMAX_FPS = 2.0floor due to the fixedmax_framesbudget spread over long durations. --detail token-burnerremoves the frame cap entirely, extracting every scene change detected by the algorithm.--max-frames Nincreases the budget linearly, directly improving frame density for the full video.- Time ranges (
--start/--end) concentrate the fixed budget on relevant segments, increasing local sampling rates. - The warning for videos >10 minutes originates in
watch.pylines 26-34 to alert users before processing resource-intensive content.
Frequently Asked Questions
What is the default frame cap for balanced detail mode?
The "balanced" detail mode defaults to 100 frames (max_frames=100). This value is defined in config.py and represents a compromise between visual coverage and API token consumption for typical video analysis tasks.
Why is my video limited to 2 fps extraction?
The codebase enforces a MAX_FPS = 2.0 constant in frames.py as a hard ceiling to prevent excessive frame generation. The auto_fps() function calculates the theoretical optimal fps by dividing the max_frames budget by duration, but floors the result at 2.0 fps for long videos to stay within budget constraints.
Does token-burner mode have any limits?
--detail token-burner removes the max_frames restriction entirely, but still respects scene detection boundaries. It extracts every frame where scene-change criteria exceed the SCENE_MIN_FRAMES threshold, potentially generating hundreds of frames for long videos with frequent cuts.
Can I permanently change the default detail level?
Yes. Create or edit ~/.config/watch/.env and add WATCH_DETAIL=token-burner to default all watch invocations to uncapped mode. The config.py module reads this file during initialization, applying your configuration without requiring command-line flags.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →