# How the Geometric Context Transformer Handles Long Sequences Beyond the 32-Frame RoPE Training Limit

> Discover how the Geometric Context Transformer processes long video sequences beyond its RoPE training limit using streaming inference KV cache sliding window attention and dynamic keyframe selection.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-31

---

**The Geometric Context Transformer (GCT) processes arbitrarily long video sequences beyond its 32-frame Rotary Position Embedding (RoPE) training limit by combining streaming inference with a KV cache, sliding-window attention, and dynamic keyframe selection to maintain constant memory usage regardless of input length.**

The Geometric Context Transformer in the `Robbyant/lingbot-map` repository overcomes the constraints of its 32-frame training window through intelligent cache management and selective attention mechanisms. While standard transformer architectures face quadratic memory growth with sequence length, GCT implements a streaming architecture that processes video frames incrementally without recomputing past representations. This approach allows the model to handle extended videos or continuous streams while preserving the temporal coherence necessary for accurate camera pose estimation and depth prediction.

## Streaming Inference with KV Cache Management

The `GCTStream` class in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) implements a streaming architecture that maintains a **key-value (KV) cache** across transformer layers. Instead of processing the entire sequence simultaneously, the model computes attention causally, using cached representations from previous frames to inform current predictions without reloading historical data.

The cache management is handled by two core components: the **AggregatorStream** and the **CameraCausalHead**. These classes expose specialized methods for cache manipulation including `_set_skip_append` for excluding non-keyframes, `_set_defer_eviction` for temporary cache preservation, and `_rollback_last_frame` for removing stale entries when flow-based criteria are not met. The constructor initializes these components through `_build_aggregator` and `_build_camera_head`, configuring them to maintain temporal state across arbitrarily long sequences.

## Sliding-Window Attention for Bounded Memory

To prevent unbounded cache growth, GCT supports **sliding-window attention** through the `sliding_window_size` parameter passed to the aggregator constructor. When configured, the model restricts attention to a fixed-size temporal window, effectively creating a rolling buffer of recent frames.

Setting `sliding_window_size=-1` enables full causal attention over the entire cached history, while positive integers enforce strict memory bounds regardless of total input length. This mechanism is implemented in the aggregator initialization between lines 52-60 of [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py), allowing inference-time configuration of the temporal receptive field without retraining the 32-frame RoPE weights.

## Dynamic Keyframe Selection Strategies

Beyond simple windowing, GCT reduces KV cache size through intelligent **keyframe selection**, treating only strategically important frames as persistent cache entries while processing others as transient observations.

### Fixed-Interval Keyframes

The simplest strategy stores every *N*-th frame as determined by the `keyframe_interval` parameter. When processing frames between keyframes, the model invokes `_set_skip_append` to prevent cache pollution, ensuring that only designated frames consume memory. This approach provides predictable, linear cache growth proportional to sequence length divided by the interval.

### Flow-Based Adaptive Keyframes

For scenarios requiring adaptive sampling, GCT implements optical-flow-based keyframe detection in the `_compute_flow_magnitude` method. The system calculates mean optical-flow magnitude between the current frame and the last keyframe; if the motion exceeds `flow_threshold` pixels or the `max_non_keyframe_gap` maximum is reached, the frame becomes a keyframe. Otherwise, `_rollback_last_frame` removes the tentative cache entry, maintaining temporal coherence without storing redundant similar frames. This logic executes within the `inference_streaming` method between lines 46-88.

## Long Sequence Processing Pipeline

The complete workflow for handling extended sequences follows five distinct stages:

1. **Cache initialization** – `self.clean_kv_cache()` removes stale tokens before processing new sequences.
2. **Scale frame processing** – The first `num_frame_for_scale` frames receive bidirectional attention to establish metric scale references.
3. **Streaming loop** – Subsequent frames process individually, with the model either deferring eviction and rolling back (flow-based) or skipping KV appends (fixed-interval) based on keyframe strategy.
4. **Prediction aggregation** – Pose, depth, and world-point predictions accumulate per-frame and concatenate at sequence end.
5. **Memory offloading** – Optional `output_device` parameters stream predictions to CPU memory, keeping GPU VRAM available for the KV cache.

Because the cache never exceeds the configured sliding window or keyframe interval, the 32-frame RoPE training constraint applies only to cached frames, not total sequence length.

## Implementation Examples

The following patterns demonstrate practical invocation of long-sequence capabilities:

```python
import torch
from lingbot_map.models.gct_stream_window import GCTStream

# ----------------------------------------------------------------------

# Example 1: Fixed-interval keyframes (default: every frame is a keyframe)

# ----------------------------------------------------------------------

model = GCTStream(
    sliding_window_size=64,          # temporal window for KV cache

    keyframe_interval=4,             # store every 4th frame in the cache

    enable_3d_rope=True,            # keep 3-D RoPE for cached frames

)

# Dummy video: 200 frames, 3×224×224 RGB

videos = torch.rand(200, 3, 224, 224).cuda()

# Run streaming inference; the model will keep only 1/4 of the frames in KV cache

outputs = model.inference_streaming(
    images=videos,
    keyframe_interval=4,
    output_device=torch.device('cpu'),   # off-load predictions to CPU

)

print(outputs["pose_enc"].shape)   # → (1, 200, 9)

# ----------------------------------------------------------------------

# Example 2: Flow-based adaptive keyframes

# ----------------------------------------------------------------------

model = GCTStream(
    sliding_window_size=128,
    enable_3d_rope=True,
)

videos = torch.rand(500, 3, 224, 224).cuda()

# The model decides dynamically when to store a frame based on motion

outputs = model.inference_streaming(
    images=videos,
    flow_threshold=2.0,           # pixels; frames with large motion become keyframes

    max_non_keyframe_gap=30,      # force a keyframe after 30 non-keyframes

    output_device=torch.device('cpu')
)

print(outputs["depth"].shape)    # → (1, 500, 224, 224, 1)

```

Key configuration parameters include:
- `sliding_window_size` – Caps temporal attention span to prevent memory overflow.
- `keyframe_interval` or `flow_threshold` – Controls cache write frequency.
- `output_device` – Moves final predictions off-GPU for very long sequences.

## Summary

- **Streaming KV cache** – `GCTStream` maintains persistent key-value caches through `AggregatorStream` and `CameraCausalHead`, enabling causal attention without full sequence recomputation.
- **Sliding-window attention** – The `sliding_window_size` parameter enforces constant memory bounds regardless of input length, configured at inference time in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py).
- **Dynamic keyframe selection** – Fixed-interval sampling via `keyframe_interval` or motion-adaptive selection via `flow_threshold` and `_compute_flow_magnitude` minimize cache footprint.
- **Cache management utilities** – Methods like `_set_skip_append`, `_rollback_last_frame`, and `clean_kv_cache` provide fine-grained control over temporal state.
- **Arbitrary sequence length** – By restricting 3-D RoPE application to cached frames only, GCT processes sequences far exceeding its 32-frame training limit while preserving temporal consistency.

## Frequently Asked Questions

### How does GCT handle sequences longer than 32 frames without retraining?

The model treats the 32-frame RoPE limit as a constraint on the **KV cache size** rather than the input length. By employing sliding-window attention and keyframe selection, GCT ensures that only a fixed number of recent frames reside in cache at any moment, applying 3-D RoPE exclusively to these cached representations. Older frames are evicted or never stored, allowing infinite sequence processing with constant memory.

### What is the difference between fixed-interval and flow-based keyframes?

**Fixed-interval keyframes** store every *N*-th frame deterministically using the `keyframe_interval` parameter, providing predictable cache growth ideal for uniform motion scenarios. **Flow-based keyframes** use `_compute_flow_magnitude` to measure optical flow between frames; frames exceeding `flow_threshold` pixel motion become keyframes, while similar frames trigger `_rollback_last_frame` to conserve cache space. Flow-based selection adapts to scene dynamics but requires additional computation.

### Can sliding-window attention and keyframe selection be used simultaneously?

Yes. The `sliding_window_size` parameter controls the maximum temporal span for attention computation, while keyframe strategies determine which frames persist in the KV cache. When both are active, the model maintains attention over the most recent `sliding_window_size` frames, but only keyframes within that window are stored for future reference. This combination provides the strictest memory bounds for ultra-long sequences.

### Does streaming inference affect prediction accuracy compared to batch processing?

According to the `Robbyant/lingbot-map` implementation, streaming inference preserves accuracy through bidirectional processing of initial scale frames (`num_frame_for_scale`) and causal attention for subsequent frames. The `CameraCausalHead` ensures that each prediction incorporates all available historical context within the cache window. Temporal consistency is maintained through the KV cache state, though extremely long-term dependencies beyond the sliding window are truncated to manage memory.