How Streaming Inference Works in LingBot-Map: KV-Cache Optimization and Causal Attention
Streaming inference in LingBot-Map is implemented through a causal transformer architecture that reuses past attention results via a sliding-window KV-cache, processing video frames sequentially after an initial scale-frame batch to achieve constant-time per-frame latency.
LingBot-Map performs online (frame-by-frame) inference using a causal architecture that maintains constant memory complexity regardless of video length. According to the Robbyant/lingbot-map source code, the implementation centers on the GCTStream class, which orchestrates a two-phase pipeline: an initial bidirectional processing of scale frames followed by causal per-frame streaming with optional keyframe subsampling.
Core Architecture Components
The streaming system comprises four tightly integrated components that manage stateful attention across video sequences:
| Component | Source File | Primary Responsibility |
|---|---|---|
| GCTStream | lingbot_map/models/gct_stream.py |
Top-level wrapper that builds the streaming-aware aggregator and causal camera head; implements the two-phase inference loop. |
| AggregatorStream | lingbot_map/aggregator/stream.py |
Transformer backbone with FlashInfer (paged) or SDPA KV-cache; handles cache initialization, updates, and sliding-window eviction. |
| CameraCausalHead | lingbot_map/heads/camera_head.py |
Lightweight pose-refinement head that operates causally while respecting the KV-cache state. |
| FlashInferKVCacheManager | lingbot_map/layers/flashinfer_cache.py |
Paged cache manager supporting sliding-window eviction for long sequences. |
Optional 3-D RoPE (Rotary Positional Embeddings) from lingbot_map/layers/rope.py maintains consistent temporal positioning for cached tokens across frames.
The Two-Phase Inference Loop
The inference_streaming method in GCTStream (lines 36-78) implements a strict separation between initialization and streaming phases to balance bidirectional context with causal efficiency.
Phase 1: Scale Frame Initialization
Before streaming begins, the model processes num_frame_for_scale frames jointly to establish a bidirectional representation. This scale phase runs a standard forward pass with causal_inference=True but processes multiple frames simultaneously:
# Initial scale processing (conceptual)
scale_output = self.forward(
scale_frames,
causal_inference=True,
num_frame_per_block=num_scale_frames
)
This phase creates the initial KV-cache entries that subsequent frames will attend to, providing global context without requiring full-sequence bidirectional attention.
Phase 2: Causal Per-Frame Streaming
Following initialization, the system enters the streaming loop where each remaining frame is processed individually. The implementation in inference_streaming (lines 47-66) handles keyframe detection and conditional KV-cache updates:
for i in range(scale_frames, total_frames):
is_keyframe = (keyframe_interval <= 1) or ((i - scale_frames) % keyframe_interval == 0)
if not is_keyframe:
self._set_skip_append(True) # Inhibit KV storage for non-keyframes
frame_output = self.forward(
frame_input,
num_frame_per_block=1,
causal_inference=True
)
if not is_keyframe:
self._set_skip_append(False) # Reset for next iteration
For non-keyframes, _set_skip_append(True) signals the cache to compute attention without storing new key-value pairs, reducing memory pressure while maintaining temporal coherence.
KV-Cache Implementation Details
The AggregatorStream class abstracts two backend implementations for key-value caching, selected automatically based on hardware availability.
FlashInfer vs SDPA Backends
FlashInfer (preferred) implements a paged KV-cache with explicit sliding-window eviction. When initialized via _init_kv_cache (lines 180-186), the system creates a lazy-initialized FlashInferKVCacheManager:
def _get_flashinfer_manager(self, ...):
if self.kv_cache_manager is None:
self.kv_cache_manager = FlashInferKVCacheManager(
sliding_window=self.kv_cache_sliding_window,
...
)
return self.kv_cache_manager
SDPA (fallback) uses a simple dictionary-based cache structure when FlashInfer is unavailable, providing compatibility at the cost of reduced memory efficiency.
Keyframe Optimization for Memory Efficiency
The keyframe mechanism controls cache growth via the keyframe_interval parameter. When processing non-keyframes, the _skip_append flag prevents cache pollution, effectively implementing temporal subsampling of the attention context. This maintains a bounded memory footprint proportional to kv_cache_sliding_window rather than sequence length.
Causal Attention and 3-D RoPE
Inside AggregatorStream._process_causal_stream (lines 514-533), the causal attention mechanism merges current tokens with cached KV-pairs:
tokens = self.global_blocks[global_idx](
tokens,
kv_cache=manager, # FlashInfer cache manager or SDPA dict
num_frames=num_frames,
...
)
Each global_blocks layer performs causal self-attention, ensuring that position i only attends to positions j ≤ i.
When enable_3d_rope=True, the aggregator generates temporal rotary embeddings via _get_3d_positions_streaming (lines 66-78), ensuring that cached tokens retain accurate positional information even as the sliding window advances through the video.
Practical Implementation Example
The following example demonstrates streaming inference with keyframe subsampling and CPU offloading:
import torch
from lingbot_map.models.gct_stream import GCTStream
# Initialize streaming model with sliding-window cache
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
enable_stream_inference=True,
kv_cache_sliding_window=64, # Retain 64 recent frames
kv_cache_scale_frames=8,
enable_3d_rope=True, # Enable temporal positional encoding
)
# Video tensor: [sequence, channels, height, width]
video = torch.rand(120, 3, 518, 518) # 120-frame video
# Execute streaming inference
outputs = model.inference_streaming(
video,
num_scale_frames=4, # Process first 4 frames bidirectionally
keyframe_interval=4, # Cache KV every 4th frame only
output_device=torch.device('cpu')
)
# Extract predictions
pose_encodings = outputs["pose_enc"] # [1, 120, 9] camera poses
depth_maps = outputs.get("depth") # Optional depth predictions
To reset state between videos, call model.clean_kv_cache() (lines 36-48 in stream.py), which clears both the FlashInfer manager and internal counters to prevent cross-sequence contamination.
Summary
- LingBot-Map implements streaming inference through a causal transformer backbone that reuses KV-cache entries across frames, achieving O(1) per-frame complexity regardless of video length.
- The two-phase pipeline processes initial scale frames bidirectionally, then switches to causal per-frame processing with optional keyframe subsampling to bound memory usage.
- FlashInfer provides a paged, sliding-window KV-cache backend, while SDPA offers a compatible fallback implementation.
- The
_set_skip_appendmechanism enables selective cache updates for non-keyframes, reducing memory pressure without sacrificing temporal consistency. - 3-D RoPE maintains accurate temporal positioning for cached tokens, ensuring stable attention patterns throughout long sequences.
Frequently Asked Questions
What is the difference between scale frames and streaming frames in LingBot-Map?
Scale frames are processed jointly at the beginning of inference to establish bidirectional context using the standard forward pass, while streaming frames are processed one-by-one with causal attention that only attends to past and current tokens. The scale phase initializes the KV-cache, and the streaming phase reuses it for efficient per-frame processing.
How does the KV-cache prevent memory overflow during long video sequences?
The implementation uses a sliding-window mechanism via kv_cache_sliding_window that evicts old frames when the cache exceeds the specified limit. Additionally, the keyframe_interval parameter allows the system to skip appending KV-pairs for non-keyframes, effectively subsampling the temporal history and maintaining constant memory usage.
Can LingBot-Map run streaming inference without FlashInfer installed?
Yes, the AggregatorStream automatically falls back to the SDPA (Scaled Dot-Product Attention) backend when FlashInfer is unavailable. This uses a dictionary-based cache instead of the paged FlashInfer manager, providing identical causal behavior with slightly higher memory overhead.
When should I use 3-D RoPE in streaming inference?
Enable enable_3d_rope=True when processing videos where temporal position significantly impacts geometric reasoning, as it applies rotary positional embeddings to the temporal dimension. This ensures that cached tokens maintain consistent relative positional encoding even as the sliding window advances, improving stability in long-sequence camera pose estimation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →