# Streaming Mode vs. Windowed Mode in LingBot-Map: Key Differences and When to Use Each

> Understand streaming vs windowed mode in LingBot-Map. Streaming ensures temporal continuity while windowed mode handles long videos within memory limits. Learn which to use.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-29

---

**Streaming mode maintains a single growing KV-cache for the entire sequence to ensure complete temporal continuity, while windowed mode processes independent segments with reset caches and stitches results via overlap keyframes to handle arbitrarily long videos within fixed memory bounds.**

LingBot-Map supports two distinct inference architectures that trade off between memory consumption and temporal coherence. Understanding the difference between streaming mode and windowed mode in LingBot-Map helps you select the appropriate strategy for your video length and GPU memory constraints.

## Core Architectural Differences

The fundamental distinction lies in how each mode manages the **KV-cache** (key-value cache), which stores the model's temporal memory during video processing.

### Streaming Mode: Continuous Context

Streaming mode utilizes the **streaming aggregator** (`AggregatorStream`) defined in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py). This mode initializes a single KV-cache via `_init_kv_cache` that persists throughout the entire inference run. The method `_process_causal_stream` executes FlashInfer-accelerated attention using this cache, which **grows linearly** with each processed frame.

Key characteristics of streaming mode:
- The KV-cache is **never cleared** during execution
- Frame-to-frame continuity is guaranteed through shared temporal state
- Memory consumption increases proportionally with sequence length

### Windowed Mode: Segmented Processing

Windowed mode implements a **sliding-window** layer in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) (lines 209-256). Instead of one continuous cache, it divides long videos into independent windows, creating a fresh `AggregatorStream` instance for each segment.

Key characteristics of windowed mode:
- Cache is **reset** at the start of every window
- Memory usage is bounded by `window_size` parameters
- Requires post-processing alignment to maintain pose continuity

## Memory Behavior and Performance Trade-offs

**Streaming mode** is ideal for moderate-length sequences where the KV-cache fits comfortably in GPU memory. As implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), the model's `inference_streaming` method (line 350) drives per-frame forward passes while accumulating context. However, for sequences exceeding approximately 3000 frames, the unbounded cache growth may exhaust available VRAM.

**Windowed mode**, accessible via `inference_windowed` (line 1022 in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)), caps memory usage by processing only one window at a time. The system accepts `--window_size` to define the number of keyframes per window and `--overlap_keyframes` to specify shared frames between adjacent windows for alignment.

The typical trade-off summary:
- **Streaming**: Smoothest pose trajectories, simplest pipeline, but linear memory growth
- **Windowed**: Constant memory footprint, supports arbitrarily long videos, but requires additional alignment logic

## Temporal Continuity Mechanisms

Maintaining coherent 3D reconstruction across window boundaries requires sophisticated stitching logic. When the KV-cache resets between windows, pose drift would naturally accumulate without correction.

The windowed implementation solves this through `_stitch_windows` and `_align_and_stitch_windows` (lines 740-966). These methods compute a **similarity transform** (scale, rotation, and translation) between overlapping keyframes. As noted around line 909, the system applies this transform to align adjacent window predictions while de-duplicating shared frames.

In contrast, streaming mode requires no such alignment because the same cache carries forward through `_process_causal_stream`, ensuring natural temporal continuity without post-processing.

## Command-Line Configuration

The CLI entry point in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (line 361) parses the `--mode` argument to select between these inference strategies.

**Running Streaming Mode (default):**

```bash
python demo.py \
    --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ \
    --mode streaming \
    --keyframe_interval 6

```

**Running Windowed Mode for Long Sequences:**

```bash
python demo.py \
    --model_path /path/to/checkpoint.pt \
    --video_path video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 2

```

## Programmatic API Access

For direct Python integration, instantiate the specific model classes:

**Streaming Inference:**

```python
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(...)
outputs = model.inference_streaming(images, num_frames)

```

**Windowed Inference:**

```python
from lingbot_map.models.gct_stream_window import GCTStreamWindow

model = GCTStreamWindow(...)
window_outputs = model.inference_windowed(images, num_frames)

```

## Summary

- **Streaming mode** keeps a persistent KV-cache in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), providing frame-to-frame continuity but consuming memory proportional to video length.
- **Windowed mode** processes segments independently through `GCTStreamWindow`, resetting the cache per window and using `_align_and_stitch_windows` to merge results.
- Use streaming for sequences under ~3000 frames where smooth trajectories are critical.
- Use windowed mode with `--window_size` and `--overlap_keyframes` for long videos or memory-constrained environments.
- Both modes rely on keyframe-based KV caches, but windowed mode limits cache size to the specified window parameters.

## Frequently Asked Questions

### When should I choose windowed mode over streaming mode?

Select windowed mode when processing very long sequences (exceeding approximately 3000 frames) or when GPU memory constraints make unbounded KV-cache growth problematic. According to the LingBot-Map source code, windowed mode maintains roughly constant memory usage regardless of video length, whereas streaming mode's linear memory growth may cause out-of-memory errors on extended videos.

### How does LingBot-Map maintain pose continuity across windows?

The system preserves continuity through **overlap keyframes** and geometric alignment. The `_align_and_stitch_windows` method in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) computes a similarity transform between overlapping frames and applies it to align adjacent window predictions. This compensates for the temporal discontinuity introduced by resetting the KV-cache between segments.

### What happens if I set the overlap too low in windowed mode?

Insufficient `--overlap_keyframes` may result in visible pose discontinuities at window boundaries. The alignment algorithm requires sufficient corresponding keyframes between adjacent windows to compute an accurate similarity transform. The repository recommends tuning this parameter alongside `--window_size` to ensure smooth transitions without excessive computational overhead.

### Does streaming mode support real-time processing better than windowed mode?

Streaming mode generally provides lower latency for frame-by-frame processing since it eliminates the overhead of window stitching and alignment post-processing. However, for extremely long continuous streams, memory pressure from the growing KV-cache may eventually force a switch to windowed mode or cause performance degradation due to memory management overhead.