How the skip_append Mechanism Enables Non-Keyframe Processing in Streaming Inference

The skip_append mechanism uses a boolean flag stored in KV-cache dictionaries to prevent non-keyframe KV pairs from being written to cache during streaming inference, allowing the model to attend to current frames without permanently storing them.

Lingbot-Map implements a streaming video processing pipeline that distinguishes between keyframes (stored in the KV cache for future attention) and non-keyframes (processed but discarded). The _skip_append flag controls this behavior by conditionally blocking cache writes while preserving full attention computation across all frames.

Where skip_append is Read: The Attention Layer

During attention computation, the model checks the _skip_append flag before deciding whether to persist the current frame’s key-value pair. In lingbot_map/layers/attention.py, the logic reads the flag from the cache dictionary and branches accordingly:


# attention.py (lines 24-31)

skip_append = kv_cache.get("_skip_append", False)   # read flag

if not skip_append:
    # KEYFRAME: store in cache (original behavior)

    ...
else:
    # NON-KEYFRAME: attend to [cached + current] without storing

    ...

When _skip_append is False, the KV pair is appended to the cache as usual. When True, the current KV is concatenated with cached values for the attention computation but discarded immediately afterward, ensuring non-keyframes do not accumulate in memory.

How skip_append is Set: The Toggle Method

The model provides a centralized helper method _set_skip_append defined in lingbot_map/models/gct_stream_window.py (lines 341-359) that propagates the flag to all relevant cache locations:

def _set_skip_append(self, skip: bool):
    """Toggle the `_skip_append` flag.
    When `skip=True` the next forward passes will **not** append KV to cache.
    """
    if hasattr(self.aggregator, 'kv_cache') and self.aggregator.kv_cache is not None:
        self.aggregator.kv_cache["_skip_append"] = skip
    if hasattr(self.aggregator, 'kv_cache_manager') and self.aggregator.kv_cache_manager is not None:
        self.aggregator.kv_cache_manager._skip_append = skip
    if self.camera_head is not None and hasattr(self.camera_head, 'kv_cache'):
        for cache_dict in self.camera_head.kv_cache:
            cache_dict["_skip_append"] = skip

This method ensures consistency across the aggregator’s cache, the FlashInfer KV cache manager, and each camera-head cache, preventing partial writes that could corrupt the cache state.

skip_append in Streaming Inference Loops

Fixed-Interval Keyframe Mode

In the standard streaming inference loop, the model toggles _skip_append around each non-keyframe forward pass. Before processing a non-keyframe, the flag is set to True; after the forward completes, it is cleared to restore normal caching behavior:

if not is_keyframe:
    self._set_skip_append(True)          # instruct attention to skip caching

frame_output = self.forward(...)

if not is_keyframe:
    self._set_skip_append(False)         # restore normal caching for next frame

This pattern appears in the streaming window implementations (around lines 4-6 in gct_stream_window.py) and ensures that only frames at fixed intervals pollute the sliding-window cache.

Flow-Based Keyframe Selection

For dynamic keyframe selection based on optical flow magnitude, the model uses a deferred eviction strategy. The forward pass runs with caching enabled, then the system decides whether to keep the frame. If the frame is rejected as a non-keyframe, the previously added KV is rolled back via _rollback_last_frame() (lines 65-99), which respects the _skip_append flag to maintain cache consistency.

This approach allows the model to evaluate frame importance using actual attention weights before committing to cache storage.

Memory Efficiency and Performance Benefits

The skip_append mechanism provides three critical advantages for streaming inference:

  • Bounded Memory Usage: Non-keyframes do not increase the KV cache size, keeping the sliding window strictly limited to keyframes regardless of input video length.
  • Reduced Latency: Skipping the cache-append operation eliminates a write-heavy step for frames that do not contribute to future attention computations.
  • Temporal Context Preservation: By still attending to the current KV during the forward pass (concatenated with cached values), the model maintains temporal coherence without permanent storage of redundant frames.

Practical Implementation Examples

You can manually control the skip_append mechanism for custom keyframe logic:


# Manual control for selective caching

model._set_skip_append(True)          # Do not cache the upcoming frame

output = model.forward(frame_tensor)  # Attend using current KV only

model._set_skip_append(False)         # Restore normal caching

For fixed-interval processing:


# Fixed-interval keyframe mode

for i, frame in enumerate(frames):
    is_keyframe = (i % keyframe_interval) == 0
    if not is_keyframe:
        model._set_skip_append(True)   # Skip caching for this non-keyframe

    out = model.forward(frame)
    if not is_keyframe:
        model._set_skip_append(False)  # Re-enable caching for the next frame

For flow-based selection using the high-level API:


# Flow-based keyframe selection

predictions = model.run_streaming(
    images,
    keyframe_interval=1,            # Disable fixed-interval logic

    flow_keyframe=True,             # Use flow magnitude to decide keyframes

)

# Internally handles _set_skip_append and possible roll-backs

Summary

  • The _skip_append flag lives inside KV-cache dictionaries and controls whether attention layers persist key-value pairs.
  • Attention layers in lingbot_map/layers/attention.py check this flag before appending to cache.
  • The _set_skip_append method in lingbot_map/models/gct_stream_window.py synchronizes the flag across all cache instances.
  • Fixed-interval streaming toggles the flag around non-keyframe forward passes.
  • Flow-based selection uses deferred caching with rollback to evaluate frames before permanent storage.
  • This mechanism keeps the KV cache bounded while preserving temporal context through full attention computation.

Frequently Asked Questions

What happens to the KV pair when skip_append is True?

When _skip_append is True, the attention layer uses the current frame’s key and value for the immediate attention computation (concatenated with cached history) but discards them immediately afterward. The pair is never written to the persistent KV cache, meaning future frames cannot attend back to this specific frame.

Where is the skip_append flag physically stored?

The flag is stored as a boolean value with the key "_skip_append" inside the KV-cache dictionaries. According to the Lingbot-Map source code, it appears in the aggregator’s cache (self.aggregator.kv_cache), the FlashInfer manager (kv_cache_manager._skip_append), and each camera-head cache dictionary (camera_head.kv_cache).

How does flow-based keyframe selection interact with skip_append?

In flow-based mode, the model temporarily disables the skip mechanism to cache the frame, evaluates its optical flow magnitude, then decides whether to keep it. If the frame is rejected as a non-keyframe, the model calls _rollback_last_frame() to remove the temporarily cached KV, effectively achieving the same result as if skip_append had been set to True initially.

Can I use skip_append for custom frame selection strategies?

Yes. The _set_skip_append method is exposed on the streaming model classes (GCTStreamWindow, GCTStreamWindowV2, and GCTStream), allowing you to manually toggle caching before arbitrary forward passes. This enables custom logic such as scene-change detection or motion-based filtering without modifying the core inference loop.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →