Streaming Mode vs. Windowed Mode in LingBot-Map: Key Differences and When to Use Each

Streaming mode maintains a single growing KV-cache for the entire sequence to ensure complete temporal continuity, while windowed mode processes independent segments with reset caches and stitches results via overlap keyframes to handle arbitrarily long videos within fixed memory bounds.

LingBot-Map supports two distinct inference architectures that trade off between memory consumption and temporal coherence. Understanding the difference between streaming mode and windowed mode in LingBot-Map helps you select the appropriate strategy for your video length and GPU memory constraints.

Core Architectural Differences

The fundamental distinction lies in how each mode manages the KV-cache (key-value cache), which stores the model's temporal memory during video processing.

Streaming Mode: Continuous Context

Streaming mode utilizes the streaming aggregator (AggregatorStream) defined in lingbot_map/aggregator/stream.py. This mode initializes a single KV-cache via _init_kv_cache that persists throughout the entire inference run. The method _process_causal_stream executes FlashInfer-accelerated attention using this cache, which grows linearly with each processed frame.

Key characteristics of streaming mode:

  • The KV-cache is never cleared during execution
  • Frame-to-frame continuity is guaranteed through shared temporal state
  • Memory consumption increases proportionally with sequence length

Windowed Mode: Segmented Processing

Windowed mode implements a sliding-window layer in lingbot_map/models/gct_stream_window.py (lines 209-256). Instead of one continuous cache, it divides long videos into independent windows, creating a fresh AggregatorStream instance for each segment.

Key characteristics of windowed mode:

  • Cache is reset at the start of every window
  • Memory usage is bounded by window_size parameters
  • Requires post-processing alignment to maintain pose continuity

Memory Behavior and Performance Trade-offs

Streaming mode is ideal for moderate-length sequences where the KV-cache fits comfortably in GPU memory. As implemented in lingbot_map/models/gct_stream.py, the model's inference_streaming method (line 350) drives per-frame forward passes while accumulating context. However, for sequences exceeding approximately 3000 frames, the unbounded cache growth may exhaust available VRAM.

Windowed mode, accessible via inference_windowed (line 1022 in lingbot_map/models/gct_stream_window.py), caps memory usage by processing only one window at a time. The system accepts --window_size to define the number of keyframes per window and --overlap_keyframes to specify shared frames between adjacent windows for alignment.

The typical trade-off summary:

  • Streaming: Smoothest pose trajectories, simplest pipeline, but linear memory growth
  • Windowed: Constant memory footprint, supports arbitrarily long videos, but requires additional alignment logic

Temporal Continuity Mechanisms

Maintaining coherent 3D reconstruction across window boundaries requires sophisticated stitching logic. When the KV-cache resets between windows, pose drift would naturally accumulate without correction.

The windowed implementation solves this through _stitch_windows and _align_and_stitch_windows (lines 740-966). These methods compute a similarity transform (scale, rotation, and translation) between overlapping keyframes. As noted around line 909, the system applies this transform to align adjacent window predictions while de-duplicating shared frames.

In contrast, streaming mode requires no such alignment because the same cache carries forward through _process_causal_stream, ensuring natural temporal continuity without post-processing.

Command-Line Configuration

The CLI entry point in demo.py (line 361) parses the --mode argument to select between these inference strategies.

Running Streaming Mode (default):

python demo.py \
    --model_path /path/to/checkpoint.pt \
    --image_folder /path/to/images/ \
    --mode streaming \
    --keyframe_interval 6

Running Windowed Mode for Long Sequences:

python demo.py \
    --model_path /path/to/checkpoint.pt \
    --video_path video.mp4 \
    --fps 10 \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8 \
    --keyframe_interval 2

Programmatic API Access

For direct Python integration, instantiate the specific model classes:

Streaming Inference:

from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(...)
outputs = model.inference_streaming(images, num_frames)

Windowed Inference:

from lingbot_map.models.gct_stream_window import GCTStreamWindow

model = GCTStreamWindow(...)
window_outputs = model.inference_windowed(images, num_frames)

Summary

  • Streaming mode keeps a persistent KV-cache in lingbot_map/aggregator/stream.py, providing frame-to-frame continuity but consuming memory proportional to video length.
  • Windowed mode processes segments independently through GCTStreamWindow, resetting the cache per window and using _align_and_stitch_windows to merge results.
  • Use streaming for sequences under ~3000 frames where smooth trajectories are critical.
  • Use windowed mode with --window_size and --overlap_keyframes for long videos or memory-constrained environments.
  • Both modes rely on keyframe-based KV caches, but windowed mode limits cache size to the specified window parameters.

Frequently Asked Questions

When should I choose windowed mode over streaming mode?

Select windowed mode when processing very long sequences (exceeding approximately 3000 frames) or when GPU memory constraints make unbounded KV-cache growth problematic. According to the LingBot-Map source code, windowed mode maintains roughly constant memory usage regardless of video length, whereas streaming mode's linear memory growth may cause out-of-memory errors on extended videos.

How does LingBot-Map maintain pose continuity across windows?

The system preserves continuity through overlap keyframes and geometric alignment. The _align_and_stitch_windows method in lingbot_map/models/gct_stream_window.py computes a similarity transform between overlapping frames and applies it to align adjacent window predictions. This compensates for the temporal discontinuity introduced by resetting the KV-cache between segments.

What happens if I set the overlap too low in windowed mode?

Insufficient --overlap_keyframes may result in visible pose discontinuities at window boundaries. The alignment algorithm requires sufficient corresponding keyframes between adjacent windows to compute an accurate similarity transform. The repository recommends tuning this parameter alongside --window_size to ensure smooth transitions without excessive computational overhead.

Does streaming mode support real-time processing better than windowed mode?

Streaming mode generally provides lower latency for frame-by-frame processing since it eliminates the overhead of window stitching and alignment post-processing. However, for extremely long continuous streams, memory pressure from the growing KV-cache may eventually force a switch to windowed mode or cause performance degradation due to memory management overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →