# What Is the RoPE Training Range in LingBot-Map and How Does It Affect Inference?

> Discover the RoPE training range in LingBot-Map and how it impacts inference length. Learn to adjust max_frame_num for longer outputs.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-30

---

**The RoPE training range in LingBot-Map defaults to 1024 frames (or tokens), which hard-caps the maximum inference length to the same value unless explicitly increased via the `max_frame_num` parameter.**

LingBot-Map uses Rotary Position Embeddings (RoPE) to encode spatial-temporal relationships within its transformer architecture. The RoPE training range determines the longest sequence length for which sinusoidal frequency tables are pre-computed during model initialization in the Robbyant/lingbot-map repository. Understanding this constraint is critical for deploying the model on video sequences or token streams that may exceed the default training limit.

## How RoPE Is Implemented in LingBot-Map

The RoPE mechanism in LingBot-Map dynamically builds frequency caches based on the input sequence length. In [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py), the forward pass extracts the sequence dimension and computes rotary components:

```python
max_position = positions.shape[1]                     # ← token count

cos_comp, sin_comp = self._compute_frequency_components(
        feature_dim, max_position, tokens.device, tokens.dtype)

```

The `positions` tensor originates from the `PositionGetter` class (lines 41-61 in the same file), which generates coordinate grids for camera-based inputs. During training, the maximum length of this tensor is constrained by the `max_frame_num` argument passed to the camera heads.

## Default Training Range and Maximum Inference Length

By default, LingBot-Map limits the RoPE training range to **1024 frames**. This cap is defined in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) within the `CameraCausalHead` and `CameraHead` constructors:

```python
def __init__(…, max_frame_num: int = 1024, …):
    …

```

Because the frequency tables are cached during the first forward pass based on this `max_frame_num` value, the model cannot reliably process sequences longer than 1024 frames during inference. Attempting to do so causes the embedding lookup to request indices beyond the cached tables, resulting in either an **index-out-of-range error** or silent **fallback to zero-padding**, depending on the tensor device.

## Extending the RoPE Training Range for Longer Sequences

To support longer video sequences or token streams, you must increase `max_frame_num` when instantiating the model. This forces the `_compute_frequency_components` method to build a larger frequency cache during initialization.

**Example: Default 1024-frame limit**

```python
from lingbot_map.heads.camera_head import CameraCausalHead

model = CameraCausalHead(
    dim_in=2048,
    trunk_depth=4,
    num_heads=16,
    max_frame_num=1024,          # default training range

    enable_3d_rope=True,
)

# Inference with 800 frames works safely

output = model(tokens_of_800_frames)

```

**Example: Extended 2048-frame limit**

```python
model = CameraCausalHead(
    dim_in=2048,
    trunk_depth=4,
    num_heads=16,
    max_frame_num=2048,          # enlarged training range

    enable_3d_rope=True,
)

# Inference with 1500 frames now supported

output = model(tokens_of_1500_frames)

```

## Technical Deep Dive into Frequency Cache Construction

The `RoPE` class in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) implements aggressive caching to avoid recomputing sinusoidal tables. The `_compute_frequency_components` method (lines 102-124) generates inverse frequency bases and position indices:

```python
def _compute_frequency_components(self, dim, seq_len, device, dtype):
    cache_key = (dim, seq_len, device, dtype)
    if cache_key not in self.frequency_cache:
        exponents = torch.arange(0, dim, 2, device=device).float() / dim
        inv_freq = 1.0 / (self.base_frequency ** exponents)
        positions = torch.arange(seq_len, device=device, dtype=inv_freq.dtype)
        angles = torch.einsum("i,j->ij", positions, inv_freq)
        # … cache cos/sin tables …

```

The `seq_len` parameter here directly corresponds to `max_position` from the forward pass, which is bounded by `max_frame_num` as implemented in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py) (lines 56-86). Once cached, these tables are immutable for the model instance, solidifying the maximum inference length at initialization time.

## Summary

- The RoPE training range in LingBot-Map defaults to **1024 frames**, defined by the `max_frame_num` parameter in camera heads.
- Exceeding this range during inference triggers index errors or undefined positional encodings due to fixed frequency cache sizes in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py).
- To increase the maximum inference length, instantiate `CameraCausalHead` or `CameraHead` with a higher `max_frame_num` value, which expands the `_compute_frequency_components` cache accordingly.
- The training range is determined at model initialization and remains fixed throughout the session unless the model is re-initialized with new parameters.

## Frequently Asked Questions

### What happens if I exceed the RoPE training range during inference?

If you input a sequence longer than the configured `max_frame_num` (default 1024), the model will attempt to access indices outside the pre-computed frequency cache in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py). This results in either a runtime index-out-of-range error or silent zero-padding of positional embeddings, causing severe degradation in spatial-temporal understanding.

### How do I increase the maximum sequence length in LingBot-Map?

Increase the `max_frame_num` argument when instantiating `CameraCausalHead` or `CameraHead` from [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py). Setting this to 2048 or 4096 during model creation forces the RoPE layer to build frequency tables for those lengths, enabling inference on correspondingly longer sequences.

### Is the RoPE training range configurable during training?

Yes, the training range is configurable at initialization via the `max_frame_num` parameter. However, once the model has processed its first batch and cached the frequency tables in `_compute_frequency_components`, the range becomes fixed for that model instance. Changing the range requires reinstantiating the model class with a new `max_frame_num` value.

### Which components control the RoPE training range?

Three primary components govern the range: `CameraCausalHead`/`CameraHead` in [`lingbot_map/heads/camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/camera_head.py) (defines `max_frame_num`), the `PositionGetter` class in [`lingbot_map/layers/rope.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/rope.py) (generates position tensors), and the `_compute_frequency_components` method (builds the sinusoidal cache). The `enable_3d_rope` flag in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py) activates the RoPE mechanism but does not affect the range limit itself.