How Windowed Inference Handles Sequences Exceeding 10,000 Frames in Lingbot-Map

Windowed inference processes arbitrarily long video streams by partitioning sequences into overlapping windows of keyframes, resetting the KV cache for each window, and stitching predictions using geometric alignment of shared overlap regions.

The GCTStream model in the Robbyant/lingbot-map repository implements windowed inference to overcome GPU memory constraints when processing video streams that exceed the KV-cache capacity of typical hardware. By dividing sequences into manageable chunks and implementing sophisticated cross-window alignment, the system can handle 10,000 frames or more while maintaining temporal consistency and bounded memory usage.

The Memory Challenge in Long-Form Video Processing

Long video sequences quickly exhaust GPU memory when storing key-value (KV) caches for transformer attention mechanisms. Traditional streaming approaches accumulate context indefinitely, but this becomes impossible for production-length videos spanning thousands of frames. The windowed inference strategy addresses this by bounding memory usage to a single window of keyframes rather than the entire sequence length.

Three-Stage Windowed Inference Pipeline

The core implementation resides in lingbot_map/models/gct_stream_window_v2.py, specifically within the inference_windowed method. The architecture operates through three distinct stages to ensure both memory efficiency and geometric consistency across window boundaries.

Window Partitioning with Keyframe-Based Overlap

The system slices input sequences into overlapping windows where window_size counts keyframes (frames stored in the KV cache) rather than raw frames. For sequences longer than this limit, the method calculates overlap using either actual frames (overlap_size) or keyframes (overlap_keyframes). When using keyframe-based overlap, the system converts this to actual-frame overlap respecting the keyframe_interval:

eff_overlap = max(num_scale_frames, overlap_keyframes * keyframe_interval)

This calculation ensures sufficient context for the scale phase while minimizing redundant computation source.

Per-Window Processing with Fresh KV Cache

Each window processes with a completely fresh KV cache via self.clean_kv_cache(), guaranteeing that GPU memory never grows beyond a single window's requirements source. Within each window, processing occurs in two distinct phases:

  • Scale phase: The initial num_scale_frames process together using bidirectional attention to establish robust spatial context.
  • Streaming phase: Remaining frames process causally with one-by-one iteration.

Whether a frame becomes a keyframe is determined by either regular intervals (keyframe_interval) or a dynamic optical flow heuristic (flow_threshold) that detects significant motion source.

Cross-Window Alignment and Geometric Stitching

After individual window processing completes, _align_and_stitch_windows merges predictions into a coherent global coordinate system. The alignment stage first estimates a similarity transform (scale s, rotation R, translation t) using paired keyframes in the overlap region through _pairwise_alignment and _warp_predictions source.

The _stitch_windows function then constructs a slice table per window and concatenates only the non-overlapping tails, eliminating duplicate frames while preserving the geometric relationships established during the alignment phase source.

Practical Implementation

Here is how to configure windowed inference for a 120-frame sequence with flow-based keyframe detection to minimize cache usage:

import torch
from lingbot_map.models.gct_stream_window_v2 import GCTStream

# Initialise the model

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    sliding_window_size=64,      # KV-cache eviction window

    enable_3d_rope=True,
)

# Dummy video: 120 frames, 3×480×640 RGB

images = torch.rand(1, 120, 3, 480, 640)

# Windowed inference with 16 keyframes per window and 4-keyframe overlap

preds = model.inference_windowed(
    images,
    window_size=16,
    overlap_keyframes=4,
    num_scale_frames=4,
    keyframe_interval=1,          # Regular mode (ignored when flow mode active)

    flow_threshold=2.0,           # Enable flow-based keyframe selection

    max_non_keyframe_gap=30,      # Safety fallback for static scenes

)

print(preds["pose_enc"].shape)       # → torch.Size([1, 120, 9])

print(preds["depth"].shape)          # → torch.Size([1, 120, 480, 640, 1])

For pure flow-based windowing without fixed keyframe intervals:

preds = model.inference_windowed(
    images,
    window_size=12,
    overlap_keyframes=None,   # Use default overlap calculation

    flow_threshold=1.5,       # Frames become keyframes when motion > 1.5 px

)

Summary

  • Windowed inference partitions sequences exceeding 10,000 frames into overlapping windows of keyframes to maintain constant GPU memory usage regardless of video length.
  • Each window initializes a fresh KV cache via clean_kv_cache(), preventing the memory accumulation that would otherwise limit sequence length.
  • Geometric alignment using similarity transforms (scale, rotation, translation) in _align_and_stitch_windows ensures consistent world coordinates across window boundaries.
  • Flow-based keyframe selection via flow_threshold dynamically reduces cache pressure by only retaining frames exhibiting significant motion.
  • The final output preserves temporal consistency through deduplication of overlapping regions in _stitch_windows while maintaining geometric fidelity.

Frequently Asked Questions

What is the maximum sequence length supported by windowed inference?

There is no hardcoded frame limit; windowed inference theoretically supports arbitrarily long sequences by processing them in independent chunks. The only constraints are available disk storage for final predictions and the window_size parameter controlling how many keyframes simultaneously reside in GPU memory. Sequences of 10,000 frames or more process identically to shorter videos, with memory usage remaining constant per window.

How does the system prevent temporal discontinuities at window boundaries?

The system employs overlapping windows and geometric alignment to maintain consistency. The _align_and_stitch_windows function estimates a similarity transform using shared keyframes in the overlap region, then applies _warp_predictions to align subsequent windows with the first window's coordinate frame before _stitch_windows deduplicates overlapping content.

What is the difference between overlap_size and overlap_keyframes?

overlap_size specifies the number of actual video frames to overlap between consecutive windows, while overlap_keyframes specifies the overlap in terms of cached keyframes. When using overlap_keyframes, the system automatically calculates the effective frame overlap as max(num_scale_frames, overlap_keyframes * keyframe_interval) to ensure sufficient geometric context for accurate alignment.

When should I use flow-based keyframe selection versus regular intervals?

Use flow_threshold when processing videos with variable motion—such as static shots followed by rapid camera movement—to reduce unnecessary KV-cache usage during still scenes. Use keyframe_interval for consistently moving cameras or when deterministic computational costs are prioritized over memory optimization. The flow-based mode dynamically adapts cache retention to content motion, often reducing memory pressure by 50% or more in static segments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →