How to Handle Sequences Exceeding the 320-Frame Training Range Without Pose Collapse in LingBot-Map

Use a combination of keyframe interval sampling, windowed inference with overlapping keyframes, and flow-based dynamic selection to keep the KV cache within the RoPE training horizon and prevent pose drift on long videos.

LingBot-Map was trained with rotary position embeddings (RoPE) covering up to ~320 frames (the "32 O-frame" window). When videos exceed this range, absolute temporal positions become out-of-distribution, causing pose predictions to drift and eventually collapse. The repository provides three complementary mechanisms in demo.py and lingbot_map/models/gct_stream_window.py to maintain stable poses on arbitrarily long sequences.

Understanding the 320-Frame Limit and Pose Collapse

RoPE encodes temporal positions as sinusoidal functions that the model learns during training. In LingBot-Map, the training distribution only covered frames up to ≈320 indices. When the KV cache stores more frames than this, positional indices grow beyond the trained range, breaking relative-position invariance. This causes the pose head to produce inconsistent world-to-camera transforms—a phenomenon called "pose collapse."

The collapse occurs because the model has never seen RoPE encodings for positions 321+, making it impossible to reliably estimate relative camera motion beyond the training horizon.

Three Mechanisms to Handle Long Sequences

Keyframe Interval Sampling

The simplest approach limits the number of cached tokens by storing only every N-th frame in the KV cache. This effectively resets the temporal range and prevents the cache from exceeding the RoPE limit.

In demo.py (lines 69-73), the --keyframe_interval argument controls this behavior. When omitted, the script automatically computes an optimal interval using ceil(num_frames / 320) (lines 74-80), ensuring the cache never grows beyond the trained range.

Windowed Inference

For extremely long sequences (e.g., >3,000 frames), windowed inference splits the video into overlapping segments. Each window maintains its own KV cache and a fresh RoPE base, ensuring the positional encoding never exceeds the training limit.

The GCTStream.inference_windowed method in lingbot_map/models/gct_stream_window.py (starting around line 959) implements this strategy. The window size is specified in keyframes, not raw frames, allowing coverage of far larger sequences. The effective frame coverage follows the formula detailed on lines 80-85 of the same file.

Flow-Based Dynamic Keyframe Selection

To ensure fast camera motions don't get missed when using fixed intervals, LingBot-Map uses optical-flow magnitude to trigger adaptive keyframes. When the average flow exceeds a threshold, the current frame is forced into the KV cache regardless of the interval.

This is controlled by --flow_threshold in demo.py (lines 70-71) with internal logic in inference_streaming (around lines 63-70 of gct_stream_window.py). The --max_non_keyframe_gap parameter provides a safety cap, forcing a keyframe after a specified number of non-keyframes even without motion.

  1. Calculate your keyframe interval using ceil(num_frames / 320) to keep the KV cache manageable. The demo script computes this automatically if --keyframe_interval is omitted.

  2. Switch to windowed mode for sequences exceeding 3,000 frames. Specify the window size in keyframes (e.g., 128) rather than raw frames to maximize coverage.

  3. Enable flow-based selection (--flow_threshold <value>) to guarantee keyframes during rapid camera movement, improving alignment across windows.

  4. Set overlapping keyframes (--overlap_keyframes) to ensure windows share enough context for reliable cross-window alignment, especially when --keyframe_interval > 1. The effective overlap is computed on lines 38-40 of gct_stream_window.py.

Implementation Details and KV Cache Management

The 3-D RoPE is enabled by default via --enable_3d_rope in the CLI and the enable_3d_rope flag in the model constructor. Positions regenerate for each window using the max_frame_num argument (see GCTStream.__init__ lines 41-54).

Fine-grained cache control is provided by _set_skip_append, _set_defer_eviction, and _rollback_last_frame (lines 41-80 of gct_stream_window.py). These methods enable both fixed-interval and flow-based policies, determining when frames are stored or discarded.

After processing each window, _align_and_stitch_windows (starting at line 998) aligns successive windows using depth-ratio scaling and similarity transforms. It concatenates results while de-duplicating overlapping frames, ensuring smooth pose trajectories across window boundaries.

Code Examples

Run these examples from the repository root to handle long sequences:


# Auto-select keyframe interval for streaming (default behavior)

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --image_folder path/to/long_sequence \
    --mask_sky

# Manual keyframe interval (keep 1 of every 10 frames)

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --image_folder path/to/long_sequence \
    --keyframe_interval 10

# Windowed inference for 5,000+ frames

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --video_path path/to/video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 2 \
    --overlap_keyframes 8

# Flow-based dynamic keyframes (trigger on >3px average flow)

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --image_folder path/to/fast_motion_sequence \
    --flow_threshold 3.0 \
    --max_non_keyframe_gap 30

# Full windowed pipeline with flow-based keyframes

python demo.py \
    --model_path path/to/lingbot-map.pt \
    --video_path path/to/video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 4 \
    --overlap_keyframes 8 \
    --flow_threshold 2.5

Summary

  • RoPE limit: LingBot-Map's positional embeddings work reliably only up to ~320 frames; beyond this, pose predictions collapse due to out-of-distribution indices.
  • Keyframe intervals: Use --keyframe_interval in demo.py to subsample frames and keep the KV cache within the training range.
  • Windowed mode: For very long videos, use inference_windowed in gct_stream_window.py to process overlapping segments with independent RoPE bases.
  • Flow adaptation: Set --flow_threshold to dynamically insert keyframes during fast motion, preventing misalignment.
  • Window stitching: The _align_and_stitch_windows function scales and merges window results, maintaining global consistency even on 25,000+ frame sequences.

Frequently Asked Questions

What causes pose collapse in LingBot-Map?

Pose collapse occurs when the KV cache stores more than ~320 frames, causing absolute temporal positions to exceed the RoPE training distribution. The model encounters sinusoidal encodings it never learned during training, producing inconsistent camera pose predictions that drift over time.

How do I choose the right keyframe interval?

Calculate ceil(total_frames / 320) to ensure the cache never exceeds the RoPE limit. For a 1,000-frame video, use --keyframe_interval 4 (or allow the auto-calculation in demo.py lines 74-80 to handle it). For stationary scenes, larger intervals work; for dynamic scenes, use smaller intervals or enable flow-based selection.

When should I use windowed mode versus streaming mode?

Use streaming mode for sequences under 3,000 frames where a single keyframe interval can manage the cache size. Switch to windowed mode (--mode windowed) for longer sequences or when you need guaranteed stability across very long videos. Windowed mode processes overlapping segments independently, preventing any single cache from exceeding the 320-frame limit.

What is the role of optical flow in preventing pose collapse?

Optical flow acts as a motion detector to trigger dynamic keyframes. When --flow_threshold is set, frames with high motion are always cached regardless of the fixed interval. This ensures that significant camera movements are preserved in the KV cache, preventing alignment errors that could propagate and cause localized pose drift within windows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →