# How Streaming Inference Works in LingBot-Map: KV-Cache Optimization and Causal Attention

> Discover how streaming inference in LingBot-Map achieves constant per-frame latency using causal attention and KV-cache optimization for efficient video processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-28

---

**Streaming inference in LingBot-Map is implemented through a causal transformer architecture that reuses past attention results via a sliding-window KV-cache, processing video frames sequentially after an initial scale-frame batch to achieve constant-time per-frame latency.**

LingBot-Map performs online (frame-by-frame) inference using a **causal architecture** that maintains constant memory complexity regardless of video length. According to the `Robbyant/lingbot-map` source code, the implementation centers on the `GCTStream` class, which orchestrates a two-phase pipeline: an initial bidirectional processing of scale frames followed by causal per-frame streaming with optional keyframe subsampling.

## Core Architecture Components

The streaming system comprises four tightly integrated components that manage stateful attention across video sequences:

| Component | Source File | Primary Responsibility |
|-----------|-------------|------------------------|
| **GCTStream** | [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) | Top-level wrapper that builds the streaming-aware aggregator and causal camera head; implements the two-phase inference loop. |
| **AggregatorStream** | [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) | Transformer backbone with FlashInfer (paged) or SDPA KV-cache; handles cache initialization, updates, and sliding-window eviction. |
| **CameraCausalHead** | [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) | Lightweight pose-refinement head that operates causally while respecting the KV-cache state. |
| **FlashInferKVCacheManager** | [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py) | Paged cache manager supporting sliding-window eviction for long sequences. |

Optional **3-D RoPE** (Rotary Positional Embeddings) from [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) maintains consistent temporal positioning for cached tokens across frames.

## The Two-Phase Inference Loop

The `inference_streaming` method in `GCTStream` (lines 36-78) implements a strict separation between initialization and streaming phases to balance bidirectional context with causal efficiency.

### Phase 1: Scale Frame Initialization

Before streaming begins, the model processes `num_frame_for_scale` frames jointly to establish a bidirectional representation. This scale phase runs a standard forward pass with `causal_inference=True` but processes multiple frames simultaneously:

```python

# Initial scale processing (conceptual)

scale_output = self.forward(
    scale_frames,
    causal_inference=True,
    num_frame_per_block=num_scale_frames
)

```

This phase creates the initial KV-cache entries that subsequent frames will attend to, providing global context without requiring full-sequence bidirectional attention.

### Phase 2: Causal Per-Frame Streaming

Following initialization, the system enters the streaming loop where each remaining frame is processed individually. The implementation in `inference_streaming` (lines 47-66) handles keyframe detection and conditional KV-cache updates:

```python
for i in range(scale_frames, total_frames):
    is_keyframe = (keyframe_interval <= 1) or ((i - scale_frames) % keyframe_interval == 0)
    
    if not is_keyframe:
        self._set_skip_append(True)  # Inhibit KV storage for non-keyframes

    
    frame_output = self.forward(
        frame_input,
        num_frame_per_block=1,
        causal_inference=True
    )
    
    if not is_keyframe:
        self._set_skip_append(False)  # Reset for next iteration

```

For non-keyframes, `_set_skip_append(True)` signals the cache to compute attention without storing new key-value pairs, reducing memory pressure while maintaining temporal coherence.

## KV-Cache Implementation Details

The `AggregatorStream` class abstracts two backend implementations for key-value caching, selected automatically based on hardware availability.

### FlashInfer vs SDPA Backends

**FlashInfer** (preferred) implements a paged KV-cache with explicit sliding-window eviction. When initialized via `_init_kv_cache` (lines 180-186), the system creates a lazy-initialized `FlashInferKVCacheManager`:

```python
def _get_flashinfer_manager(self, ...):
    if self.kv_cache_manager is None:
        self.kv_cache_manager = FlashInferKVCacheManager(
            sliding_window=self.kv_cache_sliding_window,
            ...
        )
    return self.kv_cache_manager

```

**SDPA** (fallback) uses a simple dictionary-based cache structure when FlashInfer is unavailable, providing compatibility at the cost of reduced memory efficiency.

### Keyframe Optimization for Memory Efficiency

The keyframe mechanism controls cache growth via the `keyframe_interval` parameter. When processing non-keyframes, the `_skip_append` flag prevents cache pollution, effectively implementing temporal subsampling of the attention context. This maintains a bounded memory footprint proportional to `kv_cache_sliding_window` rather than sequence length.

## Causal Attention and 3-D RoPE

Inside `AggregatorStream._process_causal_stream` (lines 514-533), the causal attention mechanism merges current tokens with cached KV-pairs:

```python
tokens = self.global_blocks[global_idx](
    tokens,
    kv_cache=manager,        # FlashInfer cache manager or SDPA dict

    num_frames=num_frames,
    ...
)

```

Each `global_blocks` layer performs causal self-attention, ensuring that position *i* only attends to positions *j ≤ i*.

When `enable_3d_rope=True`, the aggregator generates temporal rotary embeddings via `_get_3d_positions_streaming` (lines 66-78), ensuring that cached tokens retain accurate positional information even as the sliding window advances through the video.

## Practical Implementation Example

The following example demonstrates streaming inference with keyframe subsampling and CPU offloading:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

# Initialize streaming model with sliding-window cache

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    enable_stream_inference=True,
    kv_cache_sliding_window=64,      # Retain 64 recent frames

    kv_cache_scale_frames=8,
    enable_3d_rope=True,            # Enable temporal positional encoding

)

# Video tensor: [sequence, channels, height, width]

video = torch.rand(120, 3, 518, 518)  # 120-frame video

# Execute streaming inference

outputs = model.inference_streaming(
    video,
    num_scale_frames=4,          # Process first 4 frames bidirectionally

    keyframe_interval=4,         # Cache KV every 4th frame only

    output_device=torch.device('cpu')
)

# Extract predictions

pose_encodings = outputs["pose_enc"]       # [1, 120, 9] camera poses

depth_maps = outputs.get("depth")          # Optional depth predictions

```

To reset state between videos, call `model.clean_kv_cache()` (lines 36-48 in [`stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/stream.py)), which clears both the FlashInfer manager and internal counters to prevent cross-sequence contamination.

## Summary

- **LingBot-Map** implements streaming inference through a causal transformer backbone that reuses KV-cache entries across frames, achieving O(1) per-frame complexity regardless of video length.
- The **two-phase pipeline** processes initial scale frames bidirectionally, then switches to causal per-frame processing with optional keyframe subsampling to bound memory usage.
- **FlashInfer** provides a paged, sliding-window KV-cache backend, while **SDPA** offers a compatible fallback implementation.
- The **`_set_skip_append`** mechanism enables selective cache updates for non-keyframes, reducing memory pressure without sacrificing temporal consistency.
- **3-D RoPE** maintains accurate temporal positioning for cached tokens, ensuring stable attention patterns throughout long sequences.

## Frequently Asked Questions

### What is the difference between scale frames and streaming frames in LingBot-Map?

Scale frames are processed jointly at the beginning of inference to establish bidirectional context using the standard forward pass, while streaming frames are processed one-by-one with causal attention that only attends to past and current tokens. The scale phase initializes the KV-cache, and the streaming phase reuses it for efficient per-frame processing.

### How does the KV-cache prevent memory overflow during long video sequences?

The implementation uses a sliding-window mechanism via `kv_cache_sliding_window` that evicts old frames when the cache exceeds the specified limit. Additionally, the `keyframe_interval` parameter allows the system to skip appending KV-pairs for non-keyframes, effectively subsampling the temporal history and maintaining constant memory usage.

### Can LingBot-Map run streaming inference without FlashInfer installed?

Yes, the `AggregatorStream` automatically falls back to the SDPA (Scaled Dot-Product Attention) backend when FlashInfer is unavailable. This uses a dictionary-based cache instead of the paged FlashInfer manager, providing identical causal behavior with slightly higher memory overhead.

### When should I use 3-D RoPE in streaming inference?

Enable `enable_3d_rope=True` when processing videos where temporal position significantly impacts geometric reasoning, as it applies rotary positional embeddings to the temporal dimension. This ensures that cached tokens maintain consistent relative positional encoding even as the sliding window advances, improving stability in long-sequence camera pose estimation.