What Is the Purpose of Scale Frames in the Bidirectional Attention Phase?

Scale frames serve as an initial bootstrap block that enables global metric scale estimation and bidirectional context sharing before the model switches to causal streaming for remaining video frames.

In the Ling-Bot-Map repository, scale frames represent a critical architectural innovation that solves the monocular scale ambiguity problem while maintaining efficient long-video processing. These early video frames are processed together as a block where every frame can attend to every other frame, fundamentally不同于 the causal processing applied to the rest of the sequence.

Global Scale Estimation Through Early Frame Aggregation

The primary purpose of scale frames is to provide absolute metric scale for downstream 3D reconstruction and navigation tasks. By processing the first num_frame_for_scale frames simultaneously, the model aggregates depth cues across multiple viewpoints to resolve scale ambiguities impossible to determine from a single frame.

This multi-frame aggregation allows the network to infer the conversion factor between relative depth and absolute metric depth. The stable reference established during this initial phase propagates through the entire sequence, ensuring that later pose and depth predictions maintain consistent real-world units.

Bidirectional Context Sharing via the Scale Token

While the architecture processes the remaining video in a causal, frame-by-frame manner, the scale frames operate as a fully-connected sub-graph through a dedicated scale token. This learnable embedding, created in lingbot_map/aggregator/stream.py within the _setup_special_tokens method, is shared exclusively among the initial scale frames【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

The scale token enables true bidirectional attention during Phase 1, allowing each scale frame to "see" and refine its predictions based on every other scale frame. This creates a richer global context than unidirectional processing could provide, significantly improving depth and pose quality for the bootstrap block before the model commits to memory-efficient streaming.

Two-Phase Architecture Implementation

The codebase explicitly splits inference into two distinct phases:

Phase 1: Scale Frame Processing

In lingbot_map/models/gct_stream_window_v2.py, the model processes scale_images = images[:, :scale_frames] with num_frame_per_block = scale_frames and causal_inference=True, but critically includes the comment indicating these frames receive bidirectional attention via the scale token【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/models/gct_stream_window_v2.py#L27-L30】.

This phase establishes:

  • A globally consistent metric scale estimate
  • Shared bidirectional context among initial frames
  • A populated KV cache that preserves this scale information

Phase 2: Streaming Frame Processing

After consuming the scale frames, the model processes remaining frames one-by-one. The scale token is not appended to these later frames, preserving strict causal inference while the KV cache continues to reference the scale-aware context established in Phase 1. This design keeps memory usage constant for arbitrarily long videos while maintaining metric accuracy.

Implementing Scale Frames in Practice

Configure the bidirectional attention phase by specifying the number of scale frames when calling the streaming inference method:


# Process the first 8 frames as scale frames (default)

model = GCTStreamWindowV2(...)          # instantiate the model

images = load_video_frames(...)         # shape [B, S, 3, H, W]

outputs = model.inference_streaming(
    images,
    num_scale_frames=8,                 # <- number of scale frames

    keyframe_interval=1,                # every frame after scale frames is a keyframe

)

# Access the pose and depth predictions for the scale block

scale_pose = outputs["pose_enc"][:, :8]      # [B, 8, 9]

scale_depth = outputs["depth"][:, :8]       # [B, 8, H, W, 1]

For faster initialization with potentially reduced scale accuracy, decrease the number of scale frames:


# Using only 4 frames for faster initialization

outputs = model.inference_streaming(
    images,
    num_scale_frames=4,
    keyframe_interval=2,                # every 2nd frame after the scale block is a keyframe

)

In both configurations, the model first executes the bidirectional attention phase on the specified num_scale_frames, then transitions seamlessly to causal streaming for the remainder of the sequence.

Summary

  • Scale frames provide the only opportunity for bidirectional attention in an otherwise causal streaming architecture.
  • The dedicated scale token creates a fully-connected attention sub-graph among initial frames, implemented in lingbot_map/aggregator/stream.py.
  • Global metric scale is resolved by aggregating depth information across the initial frame block, solving monocular scale ambiguity.
  • The two-phase design (bidirectional bootstrap followed by causal streaming) balances accuracy with memory efficiency for long-video processing.
  • The parameter num_scale_frames controls the trade-off between initialization speed and scale estimation accuracy.

Frequently Asked Questions

How does the scale token differ from other special tokens in Ling-Bot-Map?

The scale token is registered alongside camera and optional register tokens in _setup_special_tokens, but unlike other tokens, it is only prepended to the initial scale frames during the bidirectional phase. According to the source code in lingbot_map/aggregator/stream.py, this selective application creates the fully-connected attention graph necessary for global scale estimation while allowing the rest of the video to maintain causal independence【/cache/repos/github.com/Robbyant/lingbot-map/main/lingbot_map/aggregator/stream.py#L64-L68】.

Why can't the model estimate scale from a single frame?

Monocular depth prediction inherently suffers from scale ambiguity—the network can predict relative depth ordering but cannot determine absolute metric distances without additional geometric constraints. By processing multiple consecutive frames as a block with bidirectional attention, Ling-Bot-Map leverages motion parallax and multi-view consistency to resolve the absolute scale factor that converts relative depths to metric measurements.

What happens if I set num_scale_frames to zero?

Setting num_scale_frames=0 would disable the bidirectional attention phase entirely, forcing the model to process all frames causally from the start. Without the initial bootstrap block, the system loses the ability to establish global metric scale, resulting in arbitrary depth scaling that drifts over time. The implementation in gct_stream_window_v2.py assumes at least a small number of scale frames to initialize the KV cache properly.

Does bidirectional attention on scale frames impact real-time performance?

The bidirectional attention phase only processes a small, fixed number of initial frames (typically 4-8), while the remaining video streams with efficient causal attention. Since the computational cost of bidirectional attention scales quadratically with sequence length but applies only to the tiny initial block, the impact on overall throughput is negligible. The design intentionally front-loads computation to the non-streaming bootstrap phase so that the causal phase can maintain consistent frame rates for arbitrarily long sequences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →