Difference Between GCTStream and GCTStreamWindow Models in Ling-Bot Map

GCTStream processes video frames with full causal attention using an unbounded KV cache, while GCTStreamWindow restricts attention to a sliding temporal window to maintain constant memory usage regardless of sequence length.

Both models extend the Ling-Bot Map (GCT) architecture for streaming inference, but they address different computational constraints. Understanding the difference between GCTStream and GCTStreamWindow models is essential for selecting the appropriate variant based on your memory budget and temporal context requirements.

Core Architectural Distinction

The primary divergence lies in how each model manages historical context during streaming inference.

Full-Causal Attention vs. Sliding Window

GCTStream implements full-causal attention, allowing each new frame to attend to every previously processed frame. This provides maximum temporal context but requires storing key-value (KV) pairs for the entire sequence.

GCTStreamWindow implements a sliding temporal window, limiting attention to only the most recent N frames (default approximately 64). Older frames are evicted from the KV cache, creating a fixed-size memory footprint.

KV Cache Management Strategies

In lingbot_map/models/gct_stream.py, the model accumulates KV pairs linearly with sequence length (O(S) complexity). This suits short videos or high-memory GPUs.

In lingbot_map/models/gct_stream_window.py, the cache respects kv_cache_sliding_window parameters, evicting older entries to maintain O(window) complexity. This makes it ideal for processing hour-long video streams on resource-constrained devices.

Memory Footprint and Performance

GCTStream allocates GPU memory proportional to the number of processed frames. While this preserves complete historical context for high-fidelity pose estimation, memory usage grows indefinitely, potentially causing out-of-memory errors on long sequences.

GCTStreamWindow caps memory usage by restricting the temporal receptive field. According to the source code in lingbot_map/models/gct_stream_window_v2.py, this variant automatically cleans old KV entries during inference_streaming calls, ensuring predictable latency even on edge devices.

Implementation Details

Both models utilize the underlying AggregatorStream class defined in lingbot_map/aggregator/stream.py, but configure it differently:

Parameter GCTStream GCTStreamWindow
sliding_window_size -1 (disabled) > 0 (e.g., 64)
kv_cache_sliding_window Not enforced Active eviction policy
kv_cache_include_scale_frames Optional Often disabled to respect constraints
Memory scaling Linear O(S) Constant O(window)

The windowed variant typically disables kv_cache_cross_frame_special to avoid storing special tokens from evicted frames, whereas the full-causal variant may retain these for comprehensive context.

Practical Code Examples

Full-Causal Streaming (GCTStream)

Use this configuration when processing short sequences where complete history improves accuracy:

from lingbot_map.models.gct_stream import GCTStream

# Full causal attention (no window limit)

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    sliding_window_size=-1,          # Default: unlimited context

    enable_3d_rope=True,
)

# Process sequential frames

preds = model.inference_streaming(frames_tensor)  # Shape: [S, 3, H, W]

print(preds["pose_enc"].shape)   # → [1, S, 9]

Sliding-Window Streaming (GCTStreamWindow)

Use this for long-duration videos where memory constraints are critical:

from lingbot_map.models.gct_stream_window import GCTStream

# Limit attention to recent 64 frames

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    sliding_window_size=64,         # Activate sliding-window mode

    enable_3d_rope=True,
)

# Process long video without memory growth

preds = model.inference_streaming(long_video_tensor)  # 2000+ frames

print(preds["pose_enc"].shape)   # → [1, 2000, 9] (cache ≤ 64 frames)

Summary

  • GCTStream maintains an unbounded KV cache for full causal attention, suitable for short videos requiring complete historical context.
  • GCTStreamWindow enforces a sliding temporal window (default ~64 frames) via kv_cache_sliding_window, providing constant memory usage for arbitrarily long sequences.
  • Both models share the same API and inference_streaming method, differing only in cache eviction policies configured through sliding_window_size.
  • The windowed variant is implemented across gct_stream_window.py and gct_stream_window_v2.py with refined eviction logic for production deployment.

Frequently Asked Questions

When should I use GCTStreamWindow over GCTStream?

Use GCTStreamWindow when processing long video streams or deploying to edge devices with limited GPU memory. The sliding window caps memory usage, preventing out-of-memory errors during hour-long recordings. Use GCTStream when processing short clips where maximum temporal context improves pose estimation accuracy, such as real-time robotics applications requiring full trajectory awareness.

Does GCTStreamWindow affect pose estimation accuracy?

Limiting the temporal window can reduce accuracy for motions requiring long-range temporal dependencies, as the model cannot attend to frames outside the window. However, for most local motion patterns, the default window size (64 frames) captures sufficient context. The trade-off favors GCTStreamWindow when memory constraints would otherwise prevent processing entirely.

Can I change the window size dynamically during inference?

No, the sliding_window_size parameter is fixed during model initialization and applied consistently through the AggregatorStream configuration. To use different window sizes, you must instantiate separate model instances. The kv_cache_sliding_window value determines the eviction threshold and remains constant throughout the inference_streaming session.

What files should I examine to understand the KV cache implementation?

Review lingbot_map/aggregator/stream.py for the core AggregatorStream logic, which handles KV-cache creation and FlashInfer integration. For window-specific eviction policies, examine lingbot_map/models/gct_stream_window_v2.py, which contains refined cache management and debugging hooks not present in the base window implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →