How Cue Timestamps Interact with Transcript-Flagged Moments in Claude Video
Cue timestamps are parsed into seconds, extracted as pinned "transcript-cue" frames, and merged with primary scene frames to create a chronologically ordered frame set that preserves user-specified moments.
The claude-video repository provides a watch skill that converts video into analyzable frames for visual question-answering. When users supply specific timestamps to flag transcript moments, the system creates cue frames that interact with automatically detected scene changes through a precise extraction and merging pipeline implemented in skills/watch/scripts/frames.py.
Understanding Frame Types in the Watch Skill
The watch skill generates two distinct frame categories that serve different analytical purposes:
- Detail frames: Captured at scene changes, uniform intervals, or keyframes with reasons like
"scene-change","uniform", or"keyframe"(extracted viaextract,extract_scene_candidates, orextract_keyframes). - Cue frames: Generated from user-supplied timestamps with
reason = "transcript-cue", created byextract_at_timestampsto capture specific transcript-flagged moments.
Parsing Cue Timestamps from User Input
When users provide timestamps like 30,1:05,90, the parse_timestamps function in skills/watch/scripts/frames.py (lines 95-103) handles conversion.
The function splits comma-separated values, trims whitespace, and converts each token to seconds using parse_time, which supports SS, MM:SS, and HH:MM:SS formats. The result is a sorted, de-duplicated list of floating-point seconds.
# From skills/watch/scripts/frames.py
def parse_timestamps(value: str | None) -> list[float]:
"""Parse a comma-separated list of times into a sorted, de-duplicated list."""
# Splits, trims, converts to seconds, sorts, and removes duplicates
This normalization ensures chronological processing regardless of input order.
Extracting Transcript-Cued Frames
The extract_at_timestamps function (lines 33-41) receives the parsed seconds and generates JPEG images for each valid timestamp within the specified video window.
Key behaviors:
- Window filtering: Timestamps outside the optional
[start, end]window are dropped and counted inmeta["dropped_out_of_window"]. - Frame limiting: If
max_framesis specified, timestamps are even-sampled (first and last preserved) before extraction. - Naming convention: Files follow the pattern
cue_XXXX.jpg. - Metadata tagging: Each frame is marked with
reason = "transcript-cue"and the metadata includesengine = "timestamps".
# From skills/watch/scripts/frames.py
def extract_at_timestamps(video_path, out_dir, timestamps, ...):
# Filters by window, samples if needed, extracts via ffmpeg
out.append({
"index": len(out),
"timestamp_seconds": t,
"path": str(path),
"reason": "transcript-cue",
})
# Returns frames and meta dictionary
return out, meta
Merging Cue Frames with Primary Frames
After extraction, merge_frames combines primary frames (scene-change/uniform/keyframe) with cue frames into a unified chronological sequence.
Critical behavior: Cue frames are pinned, meaning they are never removed during the merge process. The function concatenates both lists, sorts by timestamp_seconds, and re-indexes from 0.
# From skills/watch/scripts/frames.py
def merge_frames(primary: list[dict], pinned: list[dict]) -> list[dict]:
"""Combine two frame lists; pinned frames (transcript cues) are never dropped."""
merged = sorted([*primary, *pinned], key=lambda f: f["timestamp_seconds"])
for i, frame in enumerate(merged):
frame["index"] = i
return merged
This guarantees that transcript-flagged moments remain in the final frame set even when primary frame budgets are constrained.
Complete Workflow Example
The following demonstrates the end-to-end interaction between cue timestamps and primary frames:
from pathlib import Path
from skills.watch.scripts import frames
# Step 1: Parse user timestamps
raw_input = "30,1:05,90"
cue_seconds = frames.parse_timestamps(raw_input) # → [30.0, 65.0, 90.0]
# Step 2: Extract scene-based primary frames
primary = frames.extract_scene_candidates("video.mp4", Path("/tmp/scene"))
# Step 3: Extract cue frames at specific timestamps
cues, meta = frames.extract_at_timestamps(
video_path="video.mp4",
out_dir=Path("/tmp/cues"),
timestamps=[30.0, 65.0],
max_frames=2
)
# Step 4: Merge preserving chronological order
final_frames = frames.merge_frames(primary, cues)
# Result includes both scene-change and transcript-cue reasons
for f in final_frames:
print(f["index"], f["timestamp_seconds"], f["reason"])
Summary
- Cue timestamps are parsed from strings like
30,1:05,90into normalized seconds and de-duplicated byparse_timestampsinskills/watch/scripts/frames.py. - Transcript-cue frames are extracted via
extract_at_timestamps, filtered by time windows, and tagged withreason = "transcript-cue"and filenames matchingcue_XXXX.jpg. - Pinned protection: Cue frames are treated as pinned inputs to
merge_frames, ensuring they survive budget constraints that might affect primary frames. - Metadata propagation: The extraction engine tracks dropped timestamps in
meta["dropped_out_of_window"]and marks cue frames withengine = "timestamps"for downstream visibility.
Frequently Asked Questions
What happens if a cue timestamp falls outside the video window?
Timestamps outside the optional [start, end] window are silently dropped during extract_at_timestamps processing. The count of dropped timestamps is recorded in meta["dropped_out_of_window"] in the returned metadata dictionary.
How does the system handle duplicate timestamps?
The parse_timestamps function automatically de-duplicates input values while maintaining chronological order. Even if a user provides 30,30,45, the resulting list contains only unique seconds [30.0, 45.0].
Why are cue frames considered "pinned" during the merge process?
Cue frames represent user-intent flags for specific transcript moments, making them higher priority than automatically detected scene changes. By marking them as pinned in merge_frames, the system guarantees these frames remain in the final set regardless of frame budget limitations applied to primary extraction.
Can I limit the number of cue frames extracted?
Yes. When max_frames is specified in extract_at_timestamps, the function performs even-sampling across the timestamp list (preserving first and last positions) to reduce the total cue count before extraction begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →