How Bidirectional Scale Inference Differs from Streaming Forward Passes in LingBot-Map

Bidirectional scale inference jointly processes the first N frames with full cross-attention to establish camera scale, while streaming forward passes handle subsequent frames one-by-one with causal attention and KV caching to maintain memory efficiency.

LingBot-Map is a video-based mapping model that processes frame sequences through two distinct computational phases. Understanding the difference between the initial bidirectional scale inference and the subsequent streaming forward passes is essential for optimizing inference pipelines and interpreting model behavior on long video streams.

What Is Bidirectional Scale Inference?

The bidirectional scale inference phase operates on the first N frames of a video sequence (defaulting to 8) to resolve depth ambiguities and establish a consistent global scale. In [gct_stream_window_v2.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the inference_streaming method handles these scale frames as a single block where information flows freely between all frames in the set.

Implementation in inference_streaming

The scale phase corresponds to Phase 1 (lines 90‑92) of the inference_streaming function. During this phase:

  • All scale frames are fed to the model simultaneously as a single tensor block
  • Bidirectional attention allows every frame to attend to every other frame in the scale window
  • A dedicated scale token aggregates information across the entire frame set and distributes it back to enable joint scale estimation

This bidirectional processing creates a fully connected attention graph where the model can resolve scale ambiguity by comparing distant frames in both temporal directions.

How Streaming Forward Passes Work

After the scale phase completes, LingBot-Map transitions to streaming mode for all remaining frames. This phase appears as Phase 2 (lines 109‑124) in the same inference_streaming method.

Causal Attention and KV Cache Management

The streaming phase processes frames individually with strict temporal constraints:

  • Causal (unidirectional) attention restricts each frame to attend only to previously processed frames
  • A KV cache stores key and value tensors from past frames to avoid redundant computation
  • Optional keyframe logic allows selective cache updates based on the keyframe_interval parameter, trading accuracy for memory savings

Unlike the scale phase, no future frame information is available during streaming inference, ensuring the model operates in real-time compatible mode with O(1) memory growth per frame.

Key Differences Between the Two Phases

Scope of Interaction

Bidirectional scale inference treats the initial frame block as a fully connected graph where attention flows both forward and backward. This enables the model to establish geometric consistency through cycle consistency checks across the initial window.

Streaming forward passes enforce a strict temporal boundary. Each frame attends only to its history, mimicking the constraints of real-time video processing where future frames are not yet available.

Token Usage

The scale phase introduces a special scale token that acts as a global information aggregator. This token collects features from all scale frames and broadcasts refined scale estimates back to each frame position.

In streaming mode, the model relies on the standard KV cache mechanism without special aggregation tokens. The cache maintains rolling history without bidirectional communication, as implemented in the core forward method shared by both phases in [gct_stream.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

Memory and Compute Trade-offs

Bidirectional processing incurs quadratic attention complexity relative to the scale window size. However, because this phase is limited to a small fixed number of frames (typically 8), the computational cost remains manageable and provides significant accuracy benefits for scale initialization.

Streaming inference achieves linear memory scaling through the KV cache, making it suitable for arbitrarily long sequences. The optional keyframe strategy further reduces memory by skipping cache updates for non-keyframes.

Prediction Consistency

Both phases output identical prediction dictionaries containing pose_enc, depth, and world_points. The scale phase provides initial pose and depth estimates that subsequent streaming frames reference through the cached state, ensuring geometric consistency across the entire sequence.

Practical Example: Running Two-Phase Inference

The following example demonstrates how to invoke both phases using the public API:

import torch
from lingbot_map.models.gct_stream_window_v2 import LingBotMap

# Load model and prepare video tensor [S, 3, H, W]

model = LingBotMap().eval()
images = torch.randn(100, 3, 480, 640)  # 100-frame video

# Run two-phase inference

predictions = model.inference_streaming(
    images=images,
    num_scale_frames=8,      # Bidirectional phase: first 8 frames

    keyframe_interval=1,     # Every frame updates KV cache

)

# Access scale-phase outputs (bidirectional processing)

scale_poses = predictions["pose_enc"][:, :8]
scale_depths = predictions["depth"][:, :8]

# Access streaming-phase outputs (causal processing)

stream_poses = predictions["pose_enc"][:, 8:]
stream_depths = predictions["depth"][:, 8:]

The inference_streaming function automatically handles the transition between the bidirectional block (lines 90‑92) and the streaming loop (lines 109‑124), returning unified prediction tensors accessible via standard dictionary keys.

Summary

  • Bidirectional scale inference processes the first N frames as a block with full cross-attention to resolve depth scale, implemented in lines 90‑92 of gct_stream_window_v2.py
  • Streaming forward passes handle remaining frames individually with causal attention and KV caching (lines 109‑124), achieving O(1) memory growth
  • The scale token enables bidirectional information flow during initialization, while the KV cache maintains unidirectional history during streaming
  • Both phases return consistent prediction dictionaries (pose_enc, depth, world_points) but differ fundamentally in attention masking and memory complexity

Frequently Asked Questions

What is the default number of scale frames in LingBot-Map?

The default value is 8 frames, configurable via the num_scale_frames parameter in the inference_streaming method. According to the repository's README.md, increasing this number improves scale estimation accuracy at the cost of higher initial computation, while reducing it may lead to scale ambiguity in challenging sequences.

Why does the streaming phase use causal attention instead of bidirectional?

Causal attention maintains temporal integrity required for real-time processing and prevents information leakage from future frames. Since streaming inference must handle arbitrarily long sequences, causal masking combined with KV caching ensures linear memory scaling rather than the quadratic growth that would occur with bidirectional attention over the full video length.

How does the scale token affect memory usage compared to standard tokens?

The scale token adds a fixed memory overhead proportional only to the number of scale frames (typically 8), not the full sequence length. While it increases memory usage during the initial phase compared to pure streaming, it avoids the unbounded memory growth that would result from maintaining bidirectional attention across the entire video. The attention mask definitions in [lingbot_map/layers/block.py](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py) demonstrate how the bidirectional mask is constrained to this initial window.

Can I disable the bidirectional scale phase and run purely streaming inference?

No, the current implementation in gct_stream_window_v2.py requires at least one scale frame to initialize the KV cache and establish the coordinate system origin. However, you can set num_scale_frames=1 to minimize the bidirectional phase, though this may result in scale drift and reduced depth accuracy on subsequent frames due to the lack of multi-view geometric constraints during initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →