# Difference Between GCTStream and GCTStreamWindow Models in Ling-Bot Map

> Understand the key differences between GCTStream and GCTStreamWindow models in Ling-Bot Map. Learn how they handle temporal attention and memory efficiency for video frame processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-25

---

**GCTStream processes video frames with full causal attention using an unbounded KV cache, while GCTStreamWindow restricts attention to a sliding temporal window to maintain constant memory usage regardless of sequence length.**

Both models extend the **Ling-Bot Map** (GCT) architecture for streaming inference, but they address different computational constraints. Understanding the difference between GCTStream and GCTStreamWindow models is essential for selecting the appropriate variant based on your memory budget and temporal context requirements.

## Core Architectural Distinction

The primary divergence lies in how each model manages historical context during streaming inference.

### Full-Causal Attention vs. Sliding Window

**`GCTStream`** implements **full-causal attention**, allowing each new frame to attend to every previously processed frame. This provides maximum temporal context but requires storing key-value (KV) pairs for the entire sequence.

**`GCTStreamWindow`** implements a **sliding temporal window**, limiting attention to only the most recent *N* frames (default approximately 64). Older frames are evicted from the KV cache, creating a fixed-size memory footprint.

### KV Cache Management Strategies

In [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), the model accumulates KV pairs linearly with sequence length (O(S) complexity). This suits short videos or high-memory GPUs.

In [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py), the cache respects `kv_cache_sliding_window` parameters, evicting older entries to maintain O(window) complexity. This makes it ideal for processing hour-long video streams on resource-constrained devices.

## Memory Footprint and Performance

**`GCTStream`** allocates GPU memory proportional to the number of processed frames. While this preserves complete historical context for high-fidelity pose estimation, memory usage grows indefinitely, potentially causing out-of-memory errors on long sequences.

**`GCTStreamWindow`** caps memory usage by restricting the temporal receptive field. According to the source code in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), this variant automatically cleans old KV entries during `inference_streaming` calls, ensuring predictable latency even on edge devices.

## Implementation Details

Both models utilize the underlying **`AggregatorStream`** class defined in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), but configure it differently:

| Parameter | GCTStream | GCTStreamWindow |
|-----------|-----------|-----------------|
| `sliding_window_size` | `-1` (disabled) | `> 0` (e.g., `64`) |
| `kv_cache_sliding_window` | Not enforced | Active eviction policy |
| `kv_cache_include_scale_frames` | Optional | Often disabled to respect constraints |
| Memory scaling | Linear O(S) | Constant O(window) |

The windowed variant typically disables `kv_cache_cross_frame_special` to avoid storing special tokens from evicted frames, whereas the full-causal variant may retain these for comprehensive context.

## Practical Code Examples

### Full-Causal Streaming (GCTStream)

Use this configuration when processing short sequences where complete history improves accuracy:

```python
from lingbot_map.models.gct_stream import GCTStream

# Full causal attention (no window limit)

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    sliding_window_size=-1,          # Default: unlimited context

    enable_3d_rope=True,
)

# Process sequential frames

preds = model.inference_streaming(frames_tensor)  # Shape: [S, 3, H, W]

print(preds["pose_enc"].shape)   # → [1, S, 9]

```

### Sliding-Window Streaming (GCTStreamWindow)

Use this for long-duration videos where memory constraints are critical:

```python
from lingbot_map.models.gct_stream_window import GCTStream

# Limit attention to recent 64 frames

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    sliding_window_size=64,         # Activate sliding-window mode

    enable_3d_rope=True,
)

# Process long video without memory growth

preds = model.inference_streaming(long_video_tensor)  # 2000+ frames

print(preds["pose_enc"].shape)   # → [1, 2000, 9] (cache ≤ 64 frames)

```

## Summary

- **GCTStream** maintains an unbounded KV cache for full causal attention, suitable for short videos requiring complete historical context.
- **GCTStreamWindow** enforces a sliding temporal window (default ~64 frames) via `kv_cache_sliding_window`, providing constant memory usage for arbitrarily long sequences.
- Both models share the same API and `inference_streaming` method, differing only in cache eviction policies configured through `sliding_window_size`.
- The windowed variant is implemented across [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py) and [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) with refined eviction logic for production deployment.

## Frequently Asked Questions

### When should I use GCTStreamWindow over GCTStream?

**Use GCTStreamWindow when processing long video streams or deploying to edge devices with limited GPU memory.** The sliding window caps memory usage, preventing out-of-memory errors during hour-long recordings. Use GCTStream when processing short clips where maximum temporal context improves pose estimation accuracy, such as real-time robotics applications requiring full trajectory awareness.

### Does GCTStreamWindow affect pose estimation accuracy?

**Limiting the temporal window can reduce accuracy for motions requiring long-range temporal dependencies**, as the model cannot attend to frames outside the window. However, for most local motion patterns, the default window size (64 frames) captures sufficient context. The trade-off favors GCTStreamWindow when memory constraints would otherwise prevent processing entirely.

### Can I change the window size dynamically during inference?

**No, the `sliding_window_size` parameter is fixed during model initialization** and applied consistently through the `AggregatorStream` configuration. To use different window sizes, you must instantiate separate model instances. The `kv_cache_sliding_window` value determines the eviction threshold and remains constant throughout the `inference_streaming` session.

### What files should I examine to understand the KV cache implementation?

**Review [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) for the core `AggregatorStream` logic**, which handles KV-cache creation and FlashInfer integration. For window-specific eviction policies, examine [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), which contains refined cache management and debugging hooks not present in the base window implementation.