# What Are Scale Frames in LingBot‑Map’s Streaming Inference Pipeline?

> Learn how LingBot Map's streaming inference pipeline uses scale frames to establish scene scale before switching to frame-by-frame inferencing for efficient processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-29

---

**Scale frames are a short bootstrap segment that the model processes together using bidirectional attention to establish scene scale before switching to causal, frame‑by‑frame streaming.**

In the LingBot‑Map repository (Robbyant/lingbot‑map), scale frames serve as the critical initialization phase that stabilizes metric depth estimation during video processing. These initial frames provide the neural network with a multi‑frame context to jointly reason about relative depth, creating a reliable reference point that subsequent frames in the **streaming inference pipeline** can leverage without recomputing scale from scratch.

## Initializing the Scale Token with Bidirectional Attention

The primary purpose of scale frames is to **initialize a learnable scale token** that captures absolute scene scale through bidirectional attention. During this bootstrap phase, the first `scale_frames` (typically 8 by default) are processed simultaneously rather than sequentially.

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the `inference_streaming` method slices out the initial segment and feeds it to `self.forward` with `num_frame_for_scale=scale_frames` and `num_frame_per_block=scale_frames`. The inline comment *“bidirectional attention via scale token”* confirms that this block allows the network to attend across all initial frames to estimate relative depth and derive a stable scale factor【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L30】.

The **scale token** itself is instantiated in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) within the `_setup_special_tokens` method, where `self.scale_token` is defined as a dedicated learnable embedding【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

## Providing Reference for Later Frames via KV Cache

After initialization, scale frames continue to serve as a fixed reference for all subsequent streaming inference. The **KV cache** stores the attention context of the scale frames (or the compressed scale token), allowing later frames to attend to the established scale without recomputing it.

The parameter `kv_cache_scale_frames` (defaulting to 8) controls how many of these initial frames are retained in memory【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L24-L27】. This cached context eliminates the need for the model to re-estimate scale from single-frame observations, which would be far less stable.

## Serving as Overlap Between Windows

When running **windowed inference**, scale frames function as the overlap region between consecutive windows. In `inference_windowed`, the default overlap is set to `num_scale_frames` (unless overridden by a user-supplied value)【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L93-L96】.

This architectural choice guarantees that every new window begins with a well-conditioned scale context, preventing drift that would otherwise accumulate at window boundaries during long video sequences.

## Enabling Flow‑Based Keyframe Selection

Scale frames provide the final anchor for **optical‑flow‑driven keyframe decisions**. Immediately after the scale phase completes, the pipeline stores `last_kf_pose_enc = all_pose_enc[0][:, -1:]` as the reference pose encoding【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L49-L52】.

This stored reference represents the last keyframe before the causal streaming loop begins, giving the flow-based selection algorithm a stable pose and scale baseline for determining when to insert new keyframes based on motion magnitude.

## Practical Implementation Examples

### Configuring Scale Frames in Streaming Inference

To explicitly control the number of scale frames when running inference:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    kv_cache_scale_frames=8,          # Number of scale frames kept in KV cache

    kv_cache_sliding_window=64,
    enable_3d_rope=True,
)
model.eval().to("cuda")

# Dummy video: 30 frames, 3×518×518 RGB

video = torch.randn(1, 30, 3, 518, 518, device="cpu")

# Treat the first 8 frames as scale frames

pred = model.inference_streaming(
    video,
    num_scale_frames=8,               # Explicit scale‑frame count

    keyframe_interval=4,              # Every 4th frame after scale phase is stored

    flow_threshold=0.0,               # Disable flow‑based keyframe mode

)
print(pred["pose_enc"].shape)   # (1, 30, 9)

```

This corresponds to the implementation in `inference_streaming` that extracts the scale segment before entering the causal loop【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L35】.

### Benchmarking with Custom Scale Frame Counts

You can vary the scale frame count from the command line when benchmarking memory usage:

```bash
python scripts/benchmark_gct_memory.py \
    --width 518 \
    --patch-size 14 \
    --num-scale-frames 8 \
    --backend flashinfer \
    --sliding-window 64 \
    --keyframe-interval 4 \
    --frame-count 1000

```

The benchmark script forwards `args.num_scale_frames` to the model constructor via the `kv_cache_scale_frames` parameter【/cache/repos/github.com/Robbyant/lingbot-map/main/scripts/benchmark_gct_memory.py#L53-L56】.

## Summary

- **Scale frames** are a bootstrap sequence (typically 8 frames) processed with bidirectional attention to initialize a learnable scale token before causal streaming begins.
- The **KV cache** retains the scale frame context (`kv_cache_scale_frames`) so downstream frames can reference stable scale without recomputation.
- During **windowed inference**, scale frames serve as the overlap between consecutive windows, preventing drift at boundaries.
- The last scale frame provides the **reference pose** for flow-based keyframe selection algorithms.
- Implementation resides primarily in [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) and [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py).

## Frequently Asked Questions

### How many scale frames should I use for optimal performance?

The default value of **8 scale frames** provides a robust balance between initialization quality and memory overhead. According to the source code, this is controlled by `kv_cache_scale_frames` in the model constructor【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L24-L27】. For scenes with very distant or complex geometry, increasing this to 12–16 frames may improve scale estimation stability, though it linearly increases the KV cache memory footprint.

### What happens if I set scale frames to zero?

Setting `num_scale_frames=0` disables the bidirectional bootstrap phase, forcing the model to estimate scale from the first causal frame alone. Without the **scale token** initialized through multi-frame attention, the pipeline loses its absolute scale reference and must rely on single-frame predictions, which significantly degrades metric depth accuracy and temporal consistency.

### How do scale frames differ from the sliding window?

**Scale frames** are a fixed prefix processed once at the start with bidirectional attention, while the **sliding window** (`kv_cache_sliding_window`) is a rolling buffer that maintains the most recent N frames during causal streaming. Scale frames establish the initial coordinate system; the sliding window limits how far back the model can attend to prevent memory growth during long videos.

### Can the scale token attend to frames outside the initial scale segment?

No. The **scale token** is computed exclusively during the bootstrap phase using the initial `scale_frames`. Once the causal streaming phase begins, subsequent frames can attend *to* the cached scale token (via KV cache), but the scale token itself does not update or attend to new frames. This design preserves a fixed scale reference throughout the entire video sequence.