GCTStream Model Architecture: Frame Blocks, Patch Embedding, and Global Blocks Explained

The GCTStream model implements a streaming-friendly vision transformer through three core components—PatchEmbed for tokenization, frame_blocks for local per-frame processing, and global_blocks for cross-frame temporal attention with KV caching.

The GCTStream model in the Robbyant/lingbot-map repository enables efficient online inference for vision tasks through a carefully designed three-stage architecture. Unlike standard vision transformers that process entire sequences offline, this architecture separates local frame processing from global temporal attention to minimize memory costs during streaming inference.

Patch Embedding Component

The PatchEmbed layer converts input images into sequences of patch tokens that the transformer can process. Located in lingbot_map/layers/patch_embed.py, this component is instantiated within the GCTBase class constructor in lingbot_map/models/gct_base.py.

The implementation uses a convolutional approach to split images into non-overlapping patches:


# From lingbot_map/layers/patch_embed.py

self.proj = nn.Conv2d(
    in_channels, 
    embed_dim, 
    kernel_size=patch_size, 
    stride=patch_size
)

Given an input image, the layer outputs a tensor of shape [B, N_patches, embed_dim] where B is the batch size and N_patches represents the total number of patches. This projection includes special tokens prepended to the sequence for classification or aggregation tasks.

Frame Blocks Component

The frame_blocks handle local per-frame processing without maintaining a KV cache. These blocks are constructed in the _build_blocks method of lingbot_map/aggregator/stream.py (lines 30-35).

Each frame block receives the same Rotary Position Embedding (RoPE) used for patch tokens:


# From lingbot_map/aggregator/stream.py lines 30-35

self.frame_blocks = nn.ModuleList([
    block_fn(**block_params, rope=self.rope) 
    for _ in range(depth)
])

Key characteristics of the frame blocks:

  • They process the current frame only, operating on patch tokens from the PatchEmbed layer
  • They do not maintain a KV cache, keeping memory overhead minimal for local processing
  • They utilize standard transformer block architectures with RoPE for position encoding

Global Blocks Component

The global_blocks enable cross-frame temporal attention through a specialized architecture that supports KV caching. Also built in _build_blocks within lingbot_map/aggregator/stream.py (lines 36-49), these blocks handle the causal streaming logic.

The implementation selects between two backend options based on the use_sdpa flag:

  • FlashInferBlock: Uses FlashInfer's paged KV-cache for optimized inference (defined in lingbot_map/layers/block.py)
  • SDPABlock: Fallback implementation using standard SDPA (Scaled Dot-Product Attention)

# From lingbot_map/aggregator/stream.py lines 36-49

GlobalBlockCls = SDPABlock if self.use_sdpa else FlashInferBlock

self.global_blocks = nn.ModuleList([
    GlobalBlockCls(
        **block_params, 
        rope=self.rope if not self.disable_global_rope else None,
        kv_cache_sliding_window=self.kv_cache_sliding_window,
        kv_cache_scale_frames=self.kv_cache_scale_frames,
        kv_cache_cross_frame_special=self.kv_cache_cross_frame_special,
        kv_cache_include_scale_frames=self.kv_cache_include_scale_frames,
        kv_cache_camera_only=self.kv_cache_camera_only
    ) 
    for _ in range(depth)
])

These blocks manage temporal dependencies across frames while controlling memory usage through configurable sliding windows and scale frame parameters.

Integration in the Streaming Aggregator

The AggregatorStream class ties these three components together to enable causal streaming inference. During the forward pass:

  1. Input images pass through PatchEmbed to generate patch tokens
  2. Tokens flow through frame_blocks for local feature extraction
  3. Processed features enter global_blocks for cross-frame attention with KV-cache updates

This separation allows the model to process new frames with only modest memory costs, as the KV cache in global blocks maintains temporal context without recomputing attention over the entire history.

Summary

  • PatchEmbed (lingbot_map/layers/patch_embed.py): Converts images to patch tokens using strided Conv2d, outputting shape [B, N_patches, embed_dim]
  • Frame Blocks (lingbot_map/aggregator/stream.py lines 30-35): Local transformer blocks processing single frames with RoPE but no KV cache
  • Global Blocks (lingbot_map/aggregator/stream.py lines 36-49): Temporal attention blocks with KV caching support via FlashInfer or SDPA backends
  • The architecture enables efficient streaming inference by separating local and global processing stages according to the lingbot-map source code

Frequently Asked Questions

What is the difference between frame_blocks and global_blocks in GCTStream?

Frame blocks process individual frames locally without maintaining a KV cache, making them memory-efficient for spatial feature extraction. Global blocks handle cross-frame temporal attention and utilize KV caching via FlashInfer or SDPA to maintain context across the video stream without recomputing attention over the full history.

How does PatchEmbed handle different image sizes in GCTStream?

The PatchEmbed layer uses a convolutional projection with kernel_size and stride both set to the patch size, dynamically splitting input images into non-overlapping patches. This approach naturally handles various image dimensions as long as they are divisible by the patch size, outputting a variable number of patches N_patches depending on the input resolution.

Why does GCTStream use FlashInfer for the global_blocks?

FlashInfer provides optimized paged KV-cache management that reduces memory overhead during streaming inference. The global_blocks optionally use FlashInfer via FlashInferBlock instead of the SDPA fallback to enable efficient attention computation over long temporal sequences while controlling memory usage through sliding windows and selective frame caching parameters.

Where is the RoPE (Rotary Position Embedding) configured in the GCTStream architecture?

RoPE is initialized at the aggregator level in lingbot_map/aggregator/stream.py and passed to both frame_blocks and global_blocks during construction. Frame blocks always receive the rope parameter, while global blocks conditionally use it based on the disable_global_rope configuration flag.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →