How Streaming Inference Works in LingBot-Map: KV-Cache Optimization and Causal Attention

Streaming inference in LingBot-Map is implemented through a causal transformer architecture that reuses past attention results via a sliding-window KV-cache, processing video frames sequentially after an initial scale-frame batch to achieve constant-time per-frame latency.

LingBot-Map performs online (frame-by-frame) inference using a causal architecture that maintains constant memory complexity regardless of video length. According to the Robbyant/lingbot-map source code, the implementation centers on the GCTStream class, which orchestrates a two-phase pipeline: an initial bidirectional processing of scale frames followed by causal per-frame streaming with optional keyframe subsampling.

Core Architecture Components

The streaming system comprises four tightly integrated components that manage stateful attention across video sequences:

Component Source File Primary Responsibility
GCTStream lingbot_map/models/gct_stream.py Top-level wrapper that builds the streaming-aware aggregator and causal camera head; implements the two-phase inference loop.
AggregatorStream lingbot_map/aggregator/stream.py Transformer backbone with FlashInfer (paged) or SDPA KV-cache; handles cache initialization, updates, and sliding-window eviction.
CameraCausalHead lingbot_map/heads/camera_head.py Lightweight pose-refinement head that operates causally while respecting the KV-cache state.
FlashInferKVCacheManager lingbot_map/layers/flashinfer_cache.py Paged cache manager supporting sliding-window eviction for long sequences.

Optional 3-D RoPE (Rotary Positional Embeddings) from lingbot_map/layers/rope.py maintains consistent temporal positioning for cached tokens across frames.

The Two-Phase Inference Loop

The inference_streaming method in GCTStream (lines 36-78) implements a strict separation between initialization and streaming phases to balance bidirectional context with causal efficiency.

Phase 1: Scale Frame Initialization

Before streaming begins, the model processes num_frame_for_scale frames jointly to establish a bidirectional representation. This scale phase runs a standard forward pass with causal_inference=True but processes multiple frames simultaneously:


# Initial scale processing (conceptual)

scale_output = self.forward(
    scale_frames,
    causal_inference=True,
    num_frame_per_block=num_scale_frames
)

This phase creates the initial KV-cache entries that subsequent frames will attend to, providing global context without requiring full-sequence bidirectional attention.

Phase 2: Causal Per-Frame Streaming

Following initialization, the system enters the streaming loop where each remaining frame is processed individually. The implementation in inference_streaming (lines 47-66) handles keyframe detection and conditional KV-cache updates:

for i in range(scale_frames, total_frames):
    is_keyframe = (keyframe_interval <= 1) or ((i - scale_frames) % keyframe_interval == 0)
    
    if not is_keyframe:
        self._set_skip_append(True)  # Inhibit KV storage for non-keyframes

    
    frame_output = self.forward(
        frame_input,
        num_frame_per_block=1,
        causal_inference=True
    )
    
    if not is_keyframe:
        self._set_skip_append(False)  # Reset for next iteration

For non-keyframes, _set_skip_append(True) signals the cache to compute attention without storing new key-value pairs, reducing memory pressure while maintaining temporal coherence.

KV-Cache Implementation Details

The AggregatorStream class abstracts two backend implementations for key-value caching, selected automatically based on hardware availability.

FlashInfer vs SDPA Backends

FlashInfer (preferred) implements a paged KV-cache with explicit sliding-window eviction. When initialized via _init_kv_cache (lines 180-186), the system creates a lazy-initialized FlashInferKVCacheManager:

def _get_flashinfer_manager(self, ...):
    if self.kv_cache_manager is None:
        self.kv_cache_manager = FlashInferKVCacheManager(
            sliding_window=self.kv_cache_sliding_window,
            ...
        )
    return self.kv_cache_manager

SDPA (fallback) uses a simple dictionary-based cache structure when FlashInfer is unavailable, providing compatibility at the cost of reduced memory efficiency.

Keyframe Optimization for Memory Efficiency

The keyframe mechanism controls cache growth via the keyframe_interval parameter. When processing non-keyframes, the _skip_append flag prevents cache pollution, effectively implementing temporal subsampling of the attention context. This maintains a bounded memory footprint proportional to kv_cache_sliding_window rather than sequence length.

Causal Attention and 3-D RoPE

Inside AggregatorStream._process_causal_stream (lines 514-533), the causal attention mechanism merges current tokens with cached KV-pairs:

tokens = self.global_blocks[global_idx](
    tokens,
    kv_cache=manager,        # FlashInfer cache manager or SDPA dict

    num_frames=num_frames,
    ...
)

Each global_blocks layer performs causal self-attention, ensuring that position i only attends to positions j ≤ i.

When enable_3d_rope=True, the aggregator generates temporal rotary embeddings via _get_3d_positions_streaming (lines 66-78), ensuring that cached tokens retain accurate positional information even as the sliding window advances through the video.

Practical Implementation Example

The following example demonstrates streaming inference with keyframe subsampling and CPU offloading:

import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize streaming model with sliding-window cache

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    enable_stream_inference=True,
    kv_cache_sliding_window=64,      # Retain 64 recent frames

    kv_cache_scale_frames=8,
    enable_3d_rope=True,            # Enable temporal positional encoding

)

# Video tensor: [sequence, channels, height, width]

video = torch.rand(120, 3, 518, 518)  # 120-frame video

# Execute streaming inference

outputs = model.inference_streaming(
    video,
    num_scale_frames=4,          # Process first 4 frames bidirectionally

    keyframe_interval=4,         # Cache KV every 4th frame only

    output_device=torch.device('cpu')
)

# Extract predictions

pose_encodings = outputs["pose_enc"]       # [1, 120, 9] camera poses

depth_maps = outputs.get("depth")          # Optional depth predictions

To reset state between videos, call model.clean_kv_cache() (lines 36-48 in stream.py), which clears both the FlashInfer manager and internal counters to prevent cross-sequence contamination.

Summary

  • LingBot-Map implements streaming inference through a causal transformer backbone that reuses KV-cache entries across frames, achieving O(1) per-frame complexity regardless of video length.
  • The two-phase pipeline processes initial scale frames bidirectionally, then switches to causal per-frame processing with optional keyframe subsampling to bound memory usage.
  • FlashInfer provides a paged, sliding-window KV-cache backend, while SDPA offers a compatible fallback implementation.
  • The _set_skip_append mechanism enables selective cache updates for non-keyframes, reducing memory pressure without sacrificing temporal consistency.
  • 3-D RoPE maintains accurate temporal positioning for cached tokens, ensuring stable attention patterns throughout long sequences.

Frequently Asked Questions

What is the difference between scale frames and streaming frames in LingBot-Map?

Scale frames are processed jointly at the beginning of inference to establish bidirectional context using the standard forward pass, while streaming frames are processed one-by-one with causal attention that only attends to past and current tokens. The scale phase initializes the KV-cache, and the streaming phase reuses it for efficient per-frame processing.

How does the KV-cache prevent memory overflow during long video sequences?

The implementation uses a sliding-window mechanism via kv_cache_sliding_window that evicts old frames when the cache exceeds the specified limit. Additionally, the keyframe_interval parameter allows the system to skip appending KV-pairs for non-keyframes, effectively subsampling the temporal history and maintaining constant memory usage.

Can LingBot-Map run streaming inference without FlashInfer installed?

Yes, the AggregatorStream automatically falls back to the SDPA (Scaled Dot-Product Attention) backend when FlashInfer is unavailable. This uses a dictionary-based cache instead of the paged FlashInfer manager, providing identical causal behavior with slightly higher memory overhead.

When should I use 3-D RoPE in streaming inference?

Enable enable_3d_rope=True when processing videos where temporal position significantly impacts geometric reasoning, as it applies rotary positional embeddings to the temporal dimension. This ensures that cached tokens maintain consistent relative positional encoding even as the sliding window advances, improving stability in long-sequence camera pose estimation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →