# How the skip_append Mechanism Enables Non-Keyframe Processing in Streaming Inference

> Learn how the skip_append mechanism optimizes streaming inference by preventing non-keyframe KV pairs from caching, allowing models to attend to current frames without permanent storage.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: internals
- Published: 2026-07-31

---

**The `skip_append` mechanism uses a boolean flag stored in KV-cache dictionaries to prevent non-keyframe KV pairs from being written to cache during streaming inference, allowing the model to attend to current frames without permanently storing them.**

Lingbot-Map implements a streaming video processing pipeline that distinguishes between **keyframes** (stored in the KV cache for future attention) and **non-keyframes** (processed but discarded). The `_skip_append` flag controls this behavior by conditionally blocking cache writes while preserving full attention computation across all frames.

## Where skip_append is Read: The Attention Layer

During attention computation, the model checks the `_skip_append` flag before deciding whether to persist the current frame’s key-value pair. In [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py), the logic reads the flag from the cache dictionary and branches accordingly:

```python

# attention.py (lines 24-31)

skip_append = kv_cache.get("_skip_append", False)   # read flag

if not skip_append:
    # KEYFRAME: store in cache (original behavior)

    ...
else:
    # NON-KEYFRAME: attend to [cached + current] without storing

    ...

```

When `_skip_append` is **False**, the KV pair is appended to the cache as usual. When **True**, the current KV is concatenated with cached values for the attention computation but discarded immediately afterward, ensuring non-keyframes do not accumulate in memory.

## How skip_append is Set: The Toggle Method

The model provides a centralized helper method `_set_skip_append` defined in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) (lines 341-359) that propagates the flag to all relevant cache locations:

```python
def _set_skip_append(self, skip: bool):
    """Toggle the `_skip_append` flag.
    When `skip=True` the next forward passes will **not** append KV to cache.
    """
    if hasattr(self.aggregator, 'kv_cache') and self.aggregator.kv_cache is not None:
        self.aggregator.kv_cache["_skip_append"] = skip
    if hasattr(self.aggregator, 'kv_cache_manager') and self.aggregator.kv_cache_manager is not None:
        self.aggregator.kv_cache_manager._skip_append = skip
    if self.camera_head is not None and hasattr(self.camera_head, 'kv_cache'):
        for cache_dict in self.camera_head.kv_cache:
            cache_dict["_skip_append"] = skip

```

This method ensures consistency across the **aggregator’s cache**, the **FlashInfer KV cache manager**, and each **camera-head cache**, preventing partial writes that could corrupt the cache state.

## skip_append in Streaming Inference Loops

### Fixed-Interval Keyframe Mode

In the standard streaming inference loop, the model toggles `_skip_append` around each non-keyframe forward pass. Before processing a non-keyframe, the flag is set to `True`; after the forward completes, it is cleared to restore normal caching behavior:

```python
if not is_keyframe:
    self._set_skip_append(True)          # instruct attention to skip caching

frame_output = self.forward(...)

if not is_keyframe:
    self._set_skip_append(False)         # restore normal caching for next frame

```

This pattern appears in the streaming window implementations (around lines 4-6 in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py)) and ensures that only frames at fixed intervals pollute the sliding-window cache.

### Flow-Based Keyframe Selection

For dynamic keyframe selection based on optical flow magnitude, the model uses a deferred eviction strategy. The forward pass runs with caching enabled, then the system decides whether to keep the frame. If the frame is rejected as a non-keyframe, the previously added KV is rolled back via `_rollback_last_frame()` (lines 65-99), which respects the `_skip_append` flag to maintain cache consistency.

This approach allows the model to evaluate frame importance using actual attention weights before committing to cache storage.

## Memory Efficiency and Performance Benefits

The `skip_append` mechanism provides three critical advantages for streaming inference:

- **Bounded Memory Usage**: Non-keyframes do not increase the KV cache size, keeping the sliding window strictly limited to keyframes regardless of input video length.
- **Reduced Latency**: Skipping the cache-append operation eliminates a write-heavy step for frames that do not contribute to future attention computations.
- **Temporal Context Preservation**: By still attending to the current KV during the forward pass (concatenated with cached values), the model maintains temporal coherence without permanent storage of redundant frames.

## Practical Implementation Examples

You can manually control the `skip_append` mechanism for custom keyframe logic:

```python

# Manual control for selective caching

model._set_skip_append(True)          # Do not cache the upcoming frame

output = model.forward(frame_tensor)  # Attend using current KV only

model._set_skip_append(False)         # Restore normal caching

```

For fixed-interval processing:

```python

# Fixed-interval keyframe mode

for i, frame in enumerate(frames):
    is_keyframe = (i % keyframe_interval) == 0
    if not is_keyframe:
        model._set_skip_append(True)   # Skip caching for this non-keyframe

    out = model.forward(frame)
    if not is_keyframe:
        model._set_skip_append(False)  # Re-enable caching for the next frame

```

For flow-based selection using the high-level API:

```python

# Flow-based keyframe selection

predictions = model.run_streaming(
    images,
    keyframe_interval=1,            # Disable fixed-interval logic

    flow_keyframe=True,             # Use flow magnitude to decide keyframes

)

# Internally handles _set_skip_append and possible roll-backs

```

## Summary

- The `_skip_append` flag lives inside KV-cache dictionaries and controls whether attention layers persist key-value pairs.
- **Attention layers** in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) check this flag before appending to cache.
- The **`_set_skip_append`** method in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) synchronizes the flag across all cache instances.
- **Fixed-interval** streaming toggles the flag around non-keyframe forward passes.
- **Flow-based** selection uses deferred caching with rollback to evaluate frames before permanent storage.
- This mechanism keeps the KV cache bounded while preserving temporal context through full attention computation.

## Frequently Asked Questions

### What happens to the KV pair when skip_append is True?

When `_skip_append` is `True`, the attention layer uses the current frame’s key and value for the immediate attention computation (concatenated with cached history) but discards them immediately afterward. The pair is never written to the persistent KV cache, meaning future frames cannot attend back to this specific frame.

### Where is the skip_append flag physically stored?

The flag is stored as a boolean value with the key `"_skip_append"` inside the KV-cache dictionaries. According to the Lingbot-Map source code, it appears in the aggregator’s cache (`self.aggregator.kv_cache`), the FlashInfer manager (`kv_cache_manager._skip_append`), and each camera-head cache dictionary (`camera_head.kv_cache`).

### How does flow-based keyframe selection interact with skip_append?

In flow-based mode, the model temporarily disables the skip mechanism to cache the frame, evaluates its optical flow magnitude, then decides whether to keep it. If the frame is rejected as a non-keyframe, the model calls `_rollback_last_frame()` to remove the temporarily cached KV, effectively achieving the same result as if `skip_append` had been set to `True` initially.

### Can I use skip_append for custom frame selection strategies?

Yes. The `_set_skip_append` method is exposed on the streaming model classes (`GCTStreamWindow`, `GCTStreamWindowV2`, and `GCTStream`), allowing you to manually toggle caching before arbitrary forward passes. This enables custom logic such as scene-change detection or motion-based filtering without modifying the core inference loop.