Claude-Video Frame Cap Limits: How Detail Modes Control Frame Budgets
TLDR: Claude-Video uses the frame_cap() function in skills/watch/scripts/config.py to enforce hard limits of 50 frames for "efficient" mode, 100 frames for "balanced" mode, and no limit for "token-burner" or "transcript" modes.
When processing video content with the bradautomates/claude-video toolkit, the frame cap limits determine how many frames the detail engine retains for analysis. These limits vary by detail mode to balance processing speed against comprehensiveness, with specific thresholds defined in the configuration layer.
Frame Cap Limits by Detail Mode
Claude-Video maps each detail mode to a specific maximum frame allowance. The system uses these caps to throttle processing intensity, ranging from aggressive downsampling to uncapped retention.
| Detail Mode | Frame Cap | Behavior |
|---|---|---|
| efficient | 50 | Uses the fast keyframe engine; caps the budget at 50 frames. |
| balanced | 100 | Uses the thorough scene-aware engine; caps the budget at 100 frames. |
| token-burner | uncapped (None) |
No hard limit—the engine may keep every frame it deems important. |
| transcript | uncapped (None) |
Replaces detail frames with transcript-derived cue frames; no frame cap applies. |
| (unknown) | 100 | Falls back to the balanced cap for unsupported modes. |
Implementation in config.py
The authoritative logic resides in skills/watch/scripts/config.py. The frame_cap() helper translates mode strings into integer limits or None for uncapped modes:
def frame_cap(detail: str) -> int | None:
if detail == "efficient":
return 50
if detail == "balanced":
return 100
if detail == "token-burner":
return None # uncapped
if detail == "transcript":
return None # uncapped
# default for unknown modes
return 100
This function is consumed by skills/watch/scripts/watch.py to constrain the detail engine's frame budget during video processing.
Validating Frame Cap Logic with Tests
The test suite in tests/test_config.py enforces the contract, ensuring that each mode resolves to its expected limit:
def test_frame_cap_mapping():
assert config.frame_cap("efficient") == 50
assert config.frame_cap("balanced") == 100
assert config.frame_cap("token-burner") is None
assert config.frame_cap("transcript") is None
assert config.frame_cap("anything-else") == 100
These assertions prevent regression when modifying frame budget policies.
Practical Usage Examples
Command-Line Interface
Specify the detail mode when invoking the watch command to control frame retention:
# Cap at 50 frames (fastest processing)
claude-video watch myvideo.mp4 --detail efficient
# Cap at 100 frames (balanced quality)
claude-video watch myvideo.mp4 --detail balanced
# No frame cap (maximum comprehension)
claude-video watch myvideo.mp4 --detail token-burner
# Transcript-only analysis (no visual frame cap)
claude-video watch myvideo.mp4 --detail transcript
Python API
Access the frame cap programmatically to validate settings before processing:
from skills.watch.scripts import config
detail = "efficient"
max_frames = config.frame_cap(detail) # Returns 50
print(f"Detail mode '{detail}' allows up to {max_frames} frames.")
Summary
- Efficient mode caps frames at 50 for rapid, low-cost processing.
- Balanced mode caps frames at 100 for standard scene-aware analysis.
- Token-burner and transcript modes are uncapped (
None), allowing unlimited frame retention or transcript-derived cues. - Unknown modes default to the 100-frame balanced cap.
- The logic resides in
skills/watch/scripts/config.pyand is validated bytests/test_config.py.
Frequently Asked Questions
What happens if I specify an unsupported detail mode?
If you pass an unknown detail mode string, the frame_cap() function returns 100 as a fallback value, treating it as balanced mode according to the default case in skills/watch/scripts/config.py.
Why is the token-burner mode uncapped?
Token-burner mode prioritizes comprehensive visual analysis over token efficiency, allowing the engine to retain every frame it deems important without the 50 or 100 frame constraints imposed on other modes.
How does transcript mode differ from other uncapped modes?
While both transcript and token-burner modes return None for frame caps, transcript mode replaces detail frames entirely with transcript-derived cue frames, whereas token-burner mode still processes visual frames without a numerical cap.
Where are the frame cap constants defined?
The frame cap values (50, 100, None) are hardcoded in the frame_cap() function within skills/watch/scripts/config.py, not in external configuration files, ensuring consistent behavior across the bradautomates/claude-video codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →