What Are Scale Frames in LingBot‑Map’s Streaming Inference Pipeline?

Scale frames are a short bootstrap segment that the model processes together using bidirectional attention to establish scene scale before switching to causal, frame‑by‑frame streaming.

In the LingBot‑Map repository (Robbyant/lingbot‑map), scale frames serve as the critical initialization phase that stabilizes metric depth estimation during video processing. These initial frames provide the neural network with a multi‑frame context to jointly reason about relative depth, creating a reliable reference point that subsequent frames in the streaming inference pipeline can leverage without recomputing scale from scratch.

Initializing the Scale Token with Bidirectional Attention

The primary purpose of scale frames is to initialize a learnable scale token that captures absolute scene scale through bidirectional attention. During this bootstrap phase, the first scale_frames (typically 8 by default) are processed simultaneously rather than sequentially.

In lingbot_map/models/gct_stream_window_v2.py, the inference_streaming method slices out the initial segment and feeds it to self.forward with num_frame_for_scale=scale_frames and num_frame_per_block=scale_frames. The inline comment “bidirectional attention via scale token” confirms that this block allows the network to attend across all initial frames to estimate relative depth and derive a stable scale factor【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L30】.

The scale token itself is instantiated in lingbot_map/aggregator/stream.py within the _setup_special_tokens method, where self.scale_token is defined as a dedicated learnable embedding【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

Providing Reference for Later Frames via KV Cache

After initialization, scale frames continue to serve as a fixed reference for all subsequent streaming inference. The KV cache stores the attention context of the scale frames (or the compressed scale token), allowing later frames to attend to the established scale without recomputing it.

The parameter kv_cache_scale_frames (defaulting to 8) controls how many of these initial frames are retained in memory【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L24-L27】. This cached context eliminates the need for the model to re-estimate scale from single-frame observations, which would be far less stable.

Serving as Overlap Between Windows

When running windowed inference, scale frames function as the overlap region between consecutive windows. In inference_windowed, the default overlap is set to num_scale_frames (unless overridden by a user-supplied value)【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L93-L96】.

This architectural choice guarantees that every new window begins with a well-conditioned scale context, preventing drift that would otherwise accumulate at window boundaries during long video sequences.

Enabling Flow‑Based Keyframe Selection

Scale frames provide the final anchor for optical‑flow‑driven keyframe decisions. Immediately after the scale phase completes, the pipeline stores last_kf_pose_enc = all_pose_enc[0][:, -1:] as the reference pose encoding【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L49-L52】.

This stored reference represents the last keyframe before the causal streaming loop begins, giving the flow-based selection algorithm a stable pose and scale baseline for determining when to insert new keyframes based on motion magnitude.

Practical Implementation Examples

Configuring Scale Frames in Streaming Inference

To explicitly control the number of scale frames when running inference:

import torch
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    kv_cache_scale_frames=8,          # Number of scale frames kept in KV cache

    kv_cache_sliding_window=64,
    enable_3d_rope=True,
)
model.eval().to("cuda")

# Dummy video: 30 frames, 3×518×518 RGB

video = torch.randn(1, 30, 3, 518, 518, device="cpu")

# Treat the first 8 frames as scale frames

pred = model.inference_streaming(
    video,
    num_scale_frames=8,               # Explicit scale‑frame count

    keyframe_interval=4,              # Every 4th frame after scale phase is stored

    flow_threshold=0.0,               # Disable flow‑based keyframe mode

)
print(pred["pose_enc"].shape)   # (1, 30, 9)

This corresponds to the implementation in inference_streaming that extracts the scale segment before entering the causal loop【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L35】.

Benchmarking with Custom Scale Frame Counts

You can vary the scale frame count from the command line when benchmarking memory usage:

python scripts/benchmark_gct_memory.py \
    --width 518 \
    --patch-size 14 \
    --num-scale-frames 8 \
    --backend flashinfer \
    --sliding-window 64 \
    --keyframe-interval 4 \
    --frame-count 1000

The benchmark script forwards args.num_scale_frames to the model constructor via the kv_cache_scale_frames parameter【/cache/repos/github.com/Robbyant/lingbot-map/main/scripts/benchmark_gct_memory.py#L53-L56】.

Summary

  • Scale frames are a bootstrap sequence (typically 8 frames) processed with bidirectional attention to initialize a learnable scale token before causal streaming begins.
  • The KV cache retains the scale frame context (kv_cache_scale_frames) so downstream frames can reference stable scale without recomputation.
  • During windowed inference, scale frames serve as the overlap between consecutive windows, preventing drift at boundaries.
  • The last scale frame provides the reference pose for flow-based keyframe selection algorithms.
  • Implementation resides primarily in lingbot_map/models/gct_stream_window_v2.py and lingbot_map/aggregator/stream.py.

Frequently Asked Questions

How many scale frames should I use for optimal performance?

The default value of 8 scale frames provides a robust balance between initialization quality and memory overhead. According to the source code, this is controlled by kv_cache_scale_frames in the model constructor【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L24-L27】. For scenes with very distant or complex geometry, increasing this to 12–16 frames may improve scale estimation stability, though it linearly increases the KV cache memory footprint.

What happens if I set scale frames to zero?

Setting num_scale_frames=0 disables the bidirectional bootstrap phase, forcing the model to estimate scale from the first causal frame alone. Without the scale token initialized through multi-frame attention, the pipeline loses its absolute scale reference and must rely on single-frame predictions, which significantly degrades metric depth accuracy and temporal consistency.

How do scale frames differ from the sliding window?

Scale frames are a fixed prefix processed once at the start with bidirectional attention, while the sliding window (kv_cache_sliding_window) is a rolling buffer that maintains the most recent N frames during causal streaming. Scale frames establish the initial coordinate system; the sliding window limits how far back the model can attend to prevent memory growth during long videos.

Can the scale token attend to frames outside the initial scale segment?

No. The scale token is computed exclusively during the bootstrap phase using the initial scale_frames. Once the causal streaming phase begins, subsequent frames can attend to the cached scale token (via KV cache), but the scale token itself does not update or attend to new frames. This design preserves a fixed scale reference throughout the entire video sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →