# What Is the Purpose of Scale Frames in the Bidirectional Attention Phase?

> Discover the purpose of scale frames in the bidirectional attention phase. Learn how they bootstrap global scale estimation and context sharing for video processing.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-31

---

**Scale frames serve as an initial bootstrap block that enables global metric scale estimation and bidirectional context sharing before the model switches to causal streaming for remaining video frames.**

In the **Ling-Bot-Map** repository, scale frames represent a critical architectural innovation that solves the monocular scale ambiguity problem while maintaining efficient long-video processing. These early video frames are processed together as a block where every frame can attend to every other frame, fundamentally不同于 the causal processing applied to the rest of the sequence.

## Global Scale Estimation Through Early Frame Aggregation

The primary purpose of scale frames is to provide **absolute metric scale** for downstream 3D reconstruction and navigation tasks. By processing the first `num_frame_for_scale` frames simultaneously, the model aggregates depth cues across multiple viewpoints to resolve scale ambiguities impossible to determine from a single frame.

This multi-frame aggregation allows the network to infer the conversion factor between relative depth and absolute metric depth. The stable reference established during this initial phase propagates through the entire sequence, ensuring that later pose and depth predictions maintain consistent real-world units.

## Bidirectional Context Sharing via the Scale Token

While the architecture processes the remaining video in a causal, frame-by-frame manner, the scale frames operate as a **fully-connected sub-graph** through a dedicated **scale token**. This learnable embedding, created in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) within the `_setup_special_tokens` method, is shared exclusively among the initial scale frames【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

The scale token enables true bidirectional attention during Phase 1, allowing each scale frame to "see" and refine its predictions based on every other scale frame. This creates a richer global context than unidirectional processing could provide, significantly improving depth and pose quality for the bootstrap block before the model commits to memory-efficient streaming.

## Two-Phase Architecture Implementation

The codebase explicitly splits inference into two distinct phases:

### Phase 1: Scale Frame Processing

In [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py), the model processes `scale_images = images[:, :scale_frames]` with `num_frame_per_block = scale_frames` and `causal_inference=True`, but critically includes the comment indicating these frames receive **bidirectional attention via the scale token**【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L30】.

This phase establishes:
- A globally consistent metric scale estimate
- Shared bidirectional context among initial frames
- A populated KV cache that preserves this scale information

### Phase 2: Streaming Frame Processing

After consuming the scale frames, the model processes remaining frames one-by-one. The scale token is **not** appended to these later frames, preserving strict causal inference while the KV cache continues to reference the scale-aware context established in Phase 1. This design keeps memory usage constant for arbitrarily long videos while maintaining metric accuracy.

## Implementing Scale Frames in Practice

Configure the bidirectional attention phase by specifying the number of scale frames when calling the streaming inference method:

```python

# Process the first 8 frames as scale frames (default)

model = GCTStreamWindowV2(...)          # instantiate the model

images = load_video_frames(...)         # shape [B, S, 3, H, W]

outputs = model.inference_streaming(
    images,
    num_scale_frames=8,                 # <- number of scale frames

    keyframe_interval=1,                # every frame after scale frames is a keyframe

)

# Access the pose and depth predictions for the scale block

scale_pose = outputs["pose_enc"][:, :8]      # [B, 8, 9]

scale_depth = outputs["depth"][:, :8]       # [B, 8, H, W, 1]

```

For faster initialization with potentially reduced scale accuracy, decrease the number of scale frames:

```python

# Using only 4 frames for faster initialization

outputs = model.inference_streaming(
    images,
    num_scale_frames=4,
    keyframe_interval=2,                # every 2nd frame after the scale block is a keyframe

)

```

In both configurations, the model first executes the **bidirectional attention phase** on the specified `num_scale_frames`, then transitions seamlessly to causal streaming for the remainder of the sequence.

## Summary

- **Scale frames** provide the only opportunity for bidirectional attention in an otherwise causal streaming architecture.
- The dedicated **scale token** creates a fully-connected attention sub-graph among initial frames, implemented in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py).
- **Global metric scale** is resolved by aggregating depth information across the initial frame block, solving monocular scale ambiguity.
- The **two-phase design** (bidirectional bootstrap followed by causal streaming) balances accuracy with memory efficiency for long-video processing.
- The parameter `num_scale_frames` controls the trade-off between initialization speed and scale estimation accuracy.

## Frequently Asked Questions

### How does the scale token differ from other special tokens in Ling-Bot-Map?

The scale token is registered alongside camera and optional register tokens in `_setup_special_tokens`, but unlike other tokens, it is **only** prepended to the initial scale frames during the bidirectional phase. According to the source code in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), this selective application creates the fully-connected attention graph necessary for global scale estimation while allowing the rest of the video to maintain causal independence【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

### Why can't the model estimate scale from a single frame?

Monocular depth prediction inherently suffers from **scale ambiguity**—the network can predict relative depth ordering but cannot determine absolute metric distances without additional geometric constraints. By processing multiple consecutive frames as a block with bidirectional attention, Ling-Bot-Map leverages motion parallax and multi-view consistency to resolve the absolute scale factor that converts relative depths to metric measurements.

### What happens if I set num_scale_frames to zero?

Setting `num_scale_frames=0` would disable the bidirectional attention phase entirely, forcing the model to process all frames causally from the start. Without the initial bootstrap block, the system loses the ability to establish global metric scale, resulting in arbitrary depth scaling that drifts over time. The implementation in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) assumes at least a small number of scale frames to initialize the KV cache properly.

### Does bidirectional attention on scale frames impact real-time performance?

The bidirectional attention phase only processes a small, fixed number of initial frames (typically 4-8), while the remaining video streams with efficient causal attention. Since the computational cost of bidirectional attention scales quadratically with sequence length but applies only to the tiny initial block, the impact on overall throughput is negligible. The design intentionally front-loads computation to the non-streaming bootstrap phase so that the causal phase can maintain consistent frame rates for arbitrarily long sequences.