Difference Between GCTStream and GCTStreamWindow Models in Ling-Bot Map
GCTStream processes video frames with full causal attention using an unbounded KV cache, while GCTStreamWindow restricts attention to a sliding temporal window to maintain constant memory usage regardless of sequence length.
Both models extend the Ling-Bot Map (GCT) architecture for streaming inference, but they address different computational constraints. Understanding the difference between GCTStream and GCTStreamWindow models is essential for selecting the appropriate variant based on your memory budget and temporal context requirements.
Core Architectural Distinction
The primary divergence lies in how each model manages historical context during streaming inference.
Full-Causal Attention vs. Sliding Window
GCTStream implements full-causal attention, allowing each new frame to attend to every previously processed frame. This provides maximum temporal context but requires storing key-value (KV) pairs for the entire sequence.
GCTStreamWindow implements a sliding temporal window, limiting attention to only the most recent N frames (default approximately 64). Older frames are evicted from the KV cache, creating a fixed-size memory footprint.
KV Cache Management Strategies
In lingbot_map/models/gct_stream.py, the model accumulates KV pairs linearly with sequence length (O(S) complexity). This suits short videos or high-memory GPUs.
In lingbot_map/models/gct_stream_window.py, the cache respects kv_cache_sliding_window parameters, evicting older entries to maintain O(window) complexity. This makes it ideal for processing hour-long video streams on resource-constrained devices.
Memory Footprint and Performance
GCTStream allocates GPU memory proportional to the number of processed frames. While this preserves complete historical context for high-fidelity pose estimation, memory usage grows indefinitely, potentially causing out-of-memory errors on long sequences.
GCTStreamWindow caps memory usage by restricting the temporal receptive field. According to the source code in lingbot_map/models/gct_stream_window_v2.py, this variant automatically cleans old KV entries during inference_streaming calls, ensuring predictable latency even on edge devices.
Implementation Details
Both models utilize the underlying AggregatorStream class defined in lingbot_map/aggregator/stream.py, but configure it differently:
| Parameter | GCTStream | GCTStreamWindow |
|---|---|---|
sliding_window_size |
-1 (disabled) |
> 0 (e.g., 64) |
kv_cache_sliding_window |
Not enforced | Active eviction policy |
kv_cache_include_scale_frames |
Optional | Often disabled to respect constraints |
| Memory scaling | Linear O(S) | Constant O(window) |
The windowed variant typically disables kv_cache_cross_frame_special to avoid storing special tokens from evicted frames, whereas the full-causal variant may retain these for comprehensive context.
Practical Code Examples
Full-Causal Streaming (GCTStream)
Use this configuration when processing short sequences where complete history improves accuracy:
from lingbot_map.models.gct_stream import GCTStream
# Full causal attention (no window limit)
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
sliding_window_size=-1, # Default: unlimited context
enable_3d_rope=True,
)
# Process sequential frames
preds = model.inference_streaming(frames_tensor) # Shape: [S, 3, H, W]
print(preds["pose_enc"].shape) # → [1, S, 9]
Sliding-Window Streaming (GCTStreamWindow)
Use this for long-duration videos where memory constraints are critical:
from lingbot_map.models.gct_stream_window import GCTStream
# Limit attention to recent 64 frames
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
sliding_window_size=64, # Activate sliding-window mode
enable_3d_rope=True,
)
# Process long video without memory growth
preds = model.inference_streaming(long_video_tensor) # 2000+ frames
print(preds["pose_enc"].shape) # → [1, 2000, 9] (cache ≤ 64 frames)
Summary
- GCTStream maintains an unbounded KV cache for full causal attention, suitable for short videos requiring complete historical context.
- GCTStreamWindow enforces a sliding temporal window (default ~64 frames) via
kv_cache_sliding_window, providing constant memory usage for arbitrarily long sequences. - Both models share the same API and
inference_streamingmethod, differing only in cache eviction policies configured throughsliding_window_size. - The windowed variant is implemented across
gct_stream_window.pyandgct_stream_window_v2.pywith refined eviction logic for production deployment.
Frequently Asked Questions
When should I use GCTStreamWindow over GCTStream?
Use GCTStreamWindow when processing long video streams or deploying to edge devices with limited GPU memory. The sliding window caps memory usage, preventing out-of-memory errors during hour-long recordings. Use GCTStream when processing short clips where maximum temporal context improves pose estimation accuracy, such as real-time robotics applications requiring full trajectory awareness.
Does GCTStreamWindow affect pose estimation accuracy?
Limiting the temporal window can reduce accuracy for motions requiring long-range temporal dependencies, as the model cannot attend to frames outside the window. However, for most local motion patterns, the default window size (64 frames) captures sufficient context. The trade-off favors GCTStreamWindow when memory constraints would otherwise prevent processing entirely.
Can I change the window size dynamically during inference?
No, the sliding_window_size parameter is fixed during model initialization and applied consistently through the AggregatorStream configuration. To use different window sizes, you must instantiate separate model instances. The kv_cache_sliding_window value determines the eviction threshold and remains constant throughout the inference_streaming session.
What files should I examine to understand the KV cache implementation?
Review lingbot_map/aggregator/stream.py for the core AggregatorStream logic, which handles KV-cache creation and FlashInfer integration. For window-specific eviction policies, examine lingbot_map/models/gct_stream_window_v2.py, which contains refined cache management and debugging hooks not present in the base window implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →