How GCTStream Implements Streaming 3D Reconstruction with KV Cache

GCTStream enables real-time 3D reconstruction by incrementally processing video frames while reusing cached attention keys and values, limiting memory growth to O(1) per frame through sliding-window eviction and selective keyframe caching.

The GCTStream model in the Robbyant/lingbot-map repository is a streaming-focused variant of the Generalized Camera Transformer (GCT) architecture. Unlike batch processing models that require entire video sequences upfront, GCTStream performs online 3D reconstruction by maintaining a KV cache (key-value cache) across frames. This design allows the model to leverage previously computed attention states while keeping memory consumption bounded for arbitrarily long videos.

Architecture Components for Streaming

The streaming capability relies on three core components that work together to manage temporal state and memory efficiently.

Inheritance from GCTBase

GCTStream inherits the core transformer backbone from GCTBase, including patch embedding layers and transformer blocks. This relationship is defined at class GCTStream(GCTBase) in lingbot_map/models/gct_stream.py. The base class provides the foundational architecture, while the streaming variant overrides specific methods to introduce KV cache management and causal processing.

AggregatorStream with Paged KV Cache

The backbone implementation uses AggregatorStream, which wraps the transformer layers with a FlashInfer-backed KV cache system. Constructed between lines 198-226 in gct_stream.py, this component:

  • Extracts multi-scale tokens from input frames
  • Stores attention keys and values in a paged cache structure
  • Implements a sliding-window eviction policy to prevent unbounded memory growth
  • Supports configurable window sizes via the kv_cache_sliding_window parameter

CameraCausalHead for Pose Refinement

The CameraCausalHead (lines 231-253) handles per-frame camera pose estimation. Unlike standard prediction heads, it receives KV-cache parameters to enable causal attention—allowing the head to attend to cached states from previous frames while maintaining temporal consistency in the reconstructed 3D geometry.

KV Cache Lifecycle Management

GCTStream exposes explicit methods to control the KV cache state across video sequences, ensuring deterministic memory behavior between different video streams.

Cache Initialization and Cleanup

The clean_kv_cache() method (lines 84-99) clears all cached keys and values before processing a new video sequence. This prevents cross-contamination between unrelated videos and resets the internal FlashInfer cache buffers to their initial state.

Selective Frame Caching with Skip Append

To manage memory for non-essential frames, the model uses _set_skip_append(skip) (lines 100-119). When invoked with True, this method prevents the current frame's KV pairs from being written to the persistent cache. This mechanism is critical for processing non-keyframe frames without exhausting GPU memory on long videos.

Cache Monitoring

The get_kv_cache_info() method (lines 120-148) returns real-time statistics about cache occupancy, memory usage, and eviction status. This allows downstream applications to monitor resource consumption and dynamically adjust processing parameters during inference.

Streaming Inference Pipeline

The inference_streaming method (lines 150-210) orchestrates the actual 3D reconstruction workflow through two distinct processing phases.

Scale-Frame Phase

During the initial scale-frame phase, the model processes a short group of frames (typically 4) together using bidirectional attention. These frames share KV cache entries because they establish the initial geometric context. The scale tokens from this phase provide a global reference for subsequent causal processing.

Frame-by-Frame Causal Processing

After establishing the scale context, the model switches to causal frame-by-frame processing. Each new frame attends to cached KV pairs from previous frames via FlashInfer's optimized attention kernels. The keyframe_interval parameter determines whether a frame's KV tensors are persisted:

  • Keyframe frames: KV pairs are appended to the cache for future reference
  • Non-keyframe frames: _set_skip_append(True) is called to discard KV pairs after attention computation

This selective caching reduces memory growth to roughly num_frames / keyframe_interval while maintaining reconstruction accuracy.

Memory Efficiency Mechanisms

Several architectural features ensure that GCTStream maintains constant memory complexity regardless of video length.

Sliding Window Eviction

The kv_cache_sliding_window parameter enforces a hard limit on cache size. When the cache exceeds the specified window size, the oldest blocks are automatically evicted. This guarantees O(1) memory usage per processed frame rather than O(N) growth with sequence length.

Special Token Persistence

With kv_cache_cross_frame_special enabled, special tokens (such as CLS tokens) are retained across evictions even when standard frame KV pairs are discarded. This preserves high-level semantic information across the entire video while evicting low-level feature details.

3D Rotary Position Embeddings

When enable_3d_rope (or enable_camera_3d_rope) is activated, the model applies rotary positional encodings (RoPE) over the temporal dimension. This maintains geometric consistency in the reconstructed 3D space while remaining compatible with KV cache reuse, as the position encodings are applied to queries and keys during cache lookup rather than invalidating cached values.

Implementation Example

The following example demonstrates instantiating GCTStream and running streaming inference on a video tensor:

import torch
from lingbot_map.models.gct_stream import GCTStream

# Build a streaming model with KV cache enabled

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    patch_embed='dinov2_vitl14_reg',
    enable_camera=True,
    enable_depth=True,
    sliding_window_size=64,          # KV cache eviction window

    kv_cache_sliding_window=64,
    kv_cache_cross_frame_special=True,
    enable_3d_rope=True,
    use_sdpa=False,                  # Use FlashInfer backend (default)

)
model.eval().cuda()

# Dummy video: 30 RGB frames of size 518×518

B, S, C, H, W = 1, 30, 3, 518, 518
dummy_frames = torch.rand(B, S, C, H, W, device='cuda')

# Run streaming inference

outputs = model.inference_streaming(
    dummy_frames,
    num_scale_frames=4,      # First 4 frames processed jointly

    keyframe_interval=2,     # Keep KV every other frame

    output_device=torch.device('cpu')
)

# Inspect results

pose_enc = outputs['pose_enc']          # [B, S, 9] camera pose encodings

depth = outputs.get('depth')            # [B, S, H, W, 1] depth maps

print('Pose shape:', pose_enc.shape)
print('KV cache stats:', model.get_kv_cache_info())

This workflow initializes the model with a 64-frame sliding window, processes 30 frames with bidirectional context for the first 4 frames, then switches to causal processing with keyframe caching every 2 frames.

Summary

  • GCTStream extends GCTBase in lingbot_map/models/gct_stream.py to enable online 3D reconstruction through incremental frame processing.
  • The AggregatorStream backbone manages a FlashInfer-backed KV cache with configurable sliding-window eviction to bound memory usage.
  • inference_streaming implements a two-phase approach: initial scale-frame processing with bidirectional attention followed by causal frame-by-frame reconstruction.
  • Selective caching via _set_skip_append and keyframe_interval parameters allows the model to skip non-essential frames, achieving O(1) memory complexity per frame.
  • Support for 3D RoPE and cross-frame special tokens ensures temporal geometric consistency without invalidating cached attention states.

Frequently Asked Questions

How does GCTStream prevent memory exhaustion on long videos?

GCTStream implements a sliding-window KV cache controlled by the kv_cache_sliding_window parameter. When the cache reaches the window limit, the oldest KV blocks are automatically evicted. Additionally, the keyframe_interval parameter ensures that only specific frames (keyframes) persist their KV pairs to the cache, while non-keyframe computations are discarded via _set_skip_append(True).

What is the difference between the scale-frame phase and frame-by-frame phase?

The scale-frame phase processes the first num_scale_frames frames simultaneously using bidirectional attention to establish initial geometric context. The frame-by-frame phase then processes each subsequent frame causally, attending only to previous frames through the KV cache. This hybrid approach balances global context establishment with efficient incremental processing.

Can GCTStream reuse the same model instance across different video sequences?

Yes, but you must call clean_kv_cache() between sequences. This method (defined at lines 84-99 in gct_stream.py) clears all cached keys and values, resets the FlashInfer cache buffers, and ensures that attention computations from the previous video do not leak into the new sequence.

What backend handles the KV cache operations?

The model uses FlashInfer as the default backend for KV cache management, as indicated by use_sdpa=False in the constructor. The low-level cache operations are implemented in lingbot_map/layers/flashinfer_cache.py, while AggregatorStream in lingbot_map/aggregator/stream.py provides the high-level interface for cache eviction and paging.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →