# How GCTStream Implements Streaming 3D Reconstruction with KV Cache

> Discover how GCTStream achieves real-time 3D reconstruction with KV cache. Learn about its O(1) memory optimization for streaming video frames and selective keyframe caching.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-25

---

**GCTStream enables real-time 3D reconstruction by incrementally processing video frames while reusing cached attention keys and values, limiting memory growth to O(1) per frame through sliding-window eviction and selective keyframe caching.**

The **GCTStream** model in the `Robbyant/lingbot-map` repository is a streaming-focused variant of the Generalized Camera Transformer (GCT) architecture. Unlike batch processing models that require entire video sequences upfront, GCTStream performs online 3D reconstruction by maintaining a **KV cache** (key-value cache) across frames. This design allows the model to leverage previously computed attention states while keeping memory consumption bounded for arbitrarily long videos.

## Architecture Components for Streaming

The streaming capability relies on three core components that work together to manage temporal state and memory efficiently.

### Inheritance from GCTBase

`GCTStream` inherits the core transformer backbone from `GCTBase`, including patch embedding layers and transformer blocks. This relationship is defined at [`class GCTStream(GCTBase)`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py#L74) in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py). The base class provides the foundational architecture, while the streaming variant overrides specific methods to introduce KV cache management and causal processing.

### AggregatorStream with Paged KV Cache

The backbone implementation uses **`AggregatorStream`**, which wraps the transformer layers with a FlashInfer-backed KV cache system. Constructed between lines 198-226 in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py), this component:

- Extracts multi-scale tokens from input frames
- Stores attention keys and values in a paged cache structure
- Implements a sliding-window eviction policy to prevent unbounded memory growth
- Supports configurable window sizes via the `kv_cache_sliding_window` parameter

### CameraCausalHead for Pose Refinement

The **`CameraCausalHead`** (lines 231-253) handles per-frame camera pose estimation. Unlike standard prediction heads, it receives KV-cache parameters to enable causal attention—allowing the head to attend to cached states from previous frames while maintaining temporal consistency in the reconstructed 3D geometry.

## KV Cache Lifecycle Management

GCTStream exposes explicit methods to control the KV cache state across video sequences, ensuring deterministic memory behavior between different video streams.

### Cache Initialization and Cleanup

The **`clean_kv_cache()`** method (lines 84-99) clears all cached keys and values before processing a new video sequence. This prevents cross-contamination between unrelated videos and resets the internal FlashInfer cache buffers to their initial state.

### Selective Frame Caching with Skip Append

To manage memory for non-essential frames, the model uses **`_set_skip_append(skip)`** (lines 100-119). When invoked with `True`, this method prevents the current frame's KV pairs from being written to the persistent cache. This mechanism is critical for processing non-keyframe frames without exhausting GPU memory on long videos.

### Cache Monitoring

The **`get_kv_cache_info()`** method (lines 120-148) returns real-time statistics about cache occupancy, memory usage, and eviction status. This allows downstream applications to monitor resource consumption and dynamically adjust processing parameters during inference.

## Streaming Inference Pipeline

The **`inference_streaming`** method (lines 150-210) orchestrates the actual 3D reconstruction workflow through two distinct processing phases.

### Scale-Frame Phase

During the initial **scale-frame phase**, the model processes a short group of frames (typically 4) together using bidirectional attention. These frames share KV cache entries because they establish the initial geometric context. The scale tokens from this phase provide a global reference for subsequent causal processing.

### Frame-by-Frame Causal Processing

After establishing the scale context, the model switches to **causal frame-by-frame processing**. Each new frame attends to cached KV pairs from previous frames via FlashInfer's optimized attention kernels. The `keyframe_interval` parameter determines whether a frame's KV tensors are persisted:

- **Keyframe frames**: KV pairs are appended to the cache for future reference
- **Non-keyframe frames**: `_set_skip_append(True)` is called to discard KV pairs after attention computation

This selective caching reduces memory growth to roughly `num_frames / keyframe_interval` while maintaining reconstruction accuracy.

## Memory Efficiency Mechanisms

Several architectural features ensure that GCTStream maintains constant memory complexity regardless of video length.

### Sliding Window Eviction

The **`kv_cache_sliding_window`** parameter enforces a hard limit on cache size. When the cache exceeds the specified window size, the oldest blocks are automatically evicted. This guarantees O(1) memory usage per processed frame rather than O(N) growth with sequence length.

### Special Token Persistence

With **`kv_cache_cross_frame_special`** enabled, special tokens (such as CLS tokens) are retained across evictions even when standard frame KV pairs are discarded. This preserves high-level semantic information across the entire video while evicting low-level feature details.

### 3D Rotary Position Embeddings

When **`enable_3d_rope`** (or `enable_camera_3d_rope`) is activated, the model applies rotary positional encodings (RoPE) over the temporal dimension. This maintains geometric consistency in the reconstructed 3D space while remaining compatible with KV cache reuse, as the position encodings are applied to queries and keys during cache lookup rather than invalidating cached values.

## Implementation Example

The following example demonstrates instantiating `GCTStream` and running streaming inference on a video tensor:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

# Build a streaming model with KV cache enabled

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    patch_embed='dinov2_vitl14_reg',
    enable_camera=True,
    enable_depth=True,
    sliding_window_size=64,          # KV cache eviction window

    kv_cache_sliding_window=64,
    kv_cache_cross_frame_special=True,
    enable_3d_rope=True,
    use_sdpa=False,                  # Use FlashInfer backend (default)

)
model.eval().cuda()

# Dummy video: 30 RGB frames of size 518×518

B, S, C, H, W = 1, 30, 3, 518, 518
dummy_frames = torch.rand(B, S, C, H, W, device='cuda')

# Run streaming inference

outputs = model.inference_streaming(
    dummy_frames,
    num_scale_frames=4,      # First 4 frames processed jointly

    keyframe_interval=2,     # Keep KV every other frame

    output_device=torch.device('cpu')
)

# Inspect results

pose_enc = outputs['pose_enc']          # [B, S, 9] camera pose encodings

depth = outputs.get('depth')            # [B, S, H, W, 1] depth maps

print('Pose shape:', pose_enc.shape)
print('KV cache stats:', model.get_kv_cache_info())

```

This workflow initializes the model with a 64-frame sliding window, processes 30 frames with bidirectional context for the first 4 frames, then switches to causal processing with keyframe caching every 2 frames.

## Summary

- **GCTStream** extends `GCTBase` in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) to enable online 3D reconstruction through incremental frame processing.
- The **`AggregatorStream`** backbone manages a FlashInfer-backed KV cache with configurable sliding-window eviction to bound memory usage.
- **`inference_streaming`** implements a two-phase approach: initial scale-frame processing with bidirectional attention followed by causal frame-by-frame reconstruction.
- Selective caching via **`_set_skip_append`** and **`keyframe_interval`** parameters allows the model to skip non-essential frames, achieving O(1) memory complexity per frame.
- Support for **3D RoPE** and cross-frame special tokens ensures temporal geometric consistency without invalidating cached attention states.

## Frequently Asked Questions

### How does GCTStream prevent memory exhaustion on long videos?

GCTStream implements a **sliding-window KV cache** controlled by the `kv_cache_sliding_window` parameter. When the cache reaches the window limit, the oldest KV blocks are automatically evicted. Additionally, the `keyframe_interval` parameter ensures that only specific frames (keyframes) persist their KV pairs to the cache, while non-keyframe computations are discarded via `_set_skip_append(True)`.

### What is the difference between the scale-frame phase and frame-by-frame phase?

The **scale-frame phase** processes the first `num_scale_frames` frames simultaneously using bidirectional attention to establish initial geometric context. The **frame-by-frame phase** then processes each subsequent frame causally, attending only to previous frames through the KV cache. This hybrid approach balances global context establishment with efficient incremental processing.

### Can GCTStream reuse the same model instance across different video sequences?

Yes, but you must call **`clean_kv_cache()`** between sequences. This method (defined at lines 84-99 in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py)) clears all cached keys and values, resets the FlashInfer cache buffers, and ensures that attention computations from the previous video do not leak into the new sequence.

### What backend handles the KV cache operations?

The model uses **FlashInfer** as the default backend for KV cache management, as indicated by `use_sdpa=False` in the constructor. The low-level cache operations are implemented in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py), while `AggregatorStream` in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) provides the high-level interface for cache eviction and paging.