# LingBot-Map Attention Backends: SDPA, FlashInfer, and Pure PyTorch Support

> Explore LingBot-Map's support for SDPA, FlashInfer, and pure PyTorch attention backends. Easily switch between optimized modes for improved performance.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-28

---

**LingBot-Map supports three interchangeable attention backends—Standard PyTorch SDPA, FlashInfer paged KV-cache attention, and pure SDPA—allowing users to switch between them via the `attention_type` configuration parameter.**

LingBot-Map is an open-source streaming transformer implementation that provides flexible attention mechanisms for efficient inference. The repository offers multiple **attention backends** to accommodate different hardware configurations and performance requirements, from standard PyTorch implementations to optimized FlashInfer kernels. Understanding these options enables developers to optimize memory usage and throughput for their specific deployment scenarios.

## Supported Attention Backends in LingBot-Map

### Standard PyTorch SDPA

The default backend utilizes PyTorch's native `scaled_dot_product_attention` through the `Attention` and `CausalAttention` classes defined in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py). This implementation requires no additional dependencies beyond standard PyTorch and supports both full-frame self-attention and causal streaming with KV-cache management. It serves as the fallback option when specialized hardware acceleration libraries are unavailable.

### FlashInfer Paged KV-Cache Attention

For GPU-accelerated inference, LingBot-Map provides `FlashInferAttention`, which wraps FlashInfer's `BatchPrefillWithPagedKVCacheWrapper` to implement memory-efficient paged attention. This backend requires the optional FlashInfer library and is optimized for FP16/BF16 operations on compatible GPUs. The paged KV-cache approach significantly reduces memory fragmentation during long-sequence streaming inference, as implemented in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py).

### Pure SDPA Attention

The `SDPAAttention` class offers a thin wrapper around PyTorch's SDPA that maintains API consistency with the other backends while isolating the attention computation. This design facilitates benchmarking comparisons and simplifies the integration of custom attention kernels, providing a clean interface that matches the hyper-parameter signatures of `Attention` and `FlashInferAttention`.

## Configuring Attention Backends

The model architecture selects backends through the `attention_type` parameter in high-level model classes like `GCTStream` and `GCTStreamWindow`. All backend classes inherit from a shared abstract interface defined in the layers module, ensuring consistent handling of hyper-parameters including number of heads, head dimension, and dropout rates.

### Default Standard Attention

```python
from lingbot_map.models.gct_stream import GCTStream

# Use standard PyTorch SDPA (default)

model = GCTStream(attention_type="standard")

```

### Enabling FlashInfer

```python
from lingbot_map.models.gct_stream import GCTStream

# Enable FlashInfer paged attention (requires flashinfer package)

model = GCTStream(attention_type="flashinfer")

```

### Pure SDPA Backend

```python
from lingbot_map.models.gct_stream import GCTStream

# Use isolated SDPA wrapper

model = GCTStream(attention_type="sdpa")

```

## Manual Backend Swapping

Advanced users can instantiate attention layers directly and replace them in existing blocks. This approach is useful for fine-grained control over specific transformer layers.

```python
from lingbot_map.layers.attention import FlashInferAttention
from lingbot_map.layers.block import Block

# Create a transformer block with default attention

block = Block(dim=512, num_heads=8)

# Replace with FlashInfer attention manually

block.attn = FlashInferAttention(
    dim=512,
    num_heads=8,
    fused_attn=True,
    enable_3d_rope=False
)

```

## Key Source Files

The attention backend implementations are organized across several core files:

- **[`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py)**: Contains the `Attention`, `CausalAttention`, `FlashInferAttention`, and `SDPAAttention` class definitions.
- **[`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py)**: Provides utilities for managing FlashInfer's paged KV-cache buffers.
- **[`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py)**: Implements the transformer `Block` class that wires the selected attention backend into the architecture.
- **[`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)**: High-level streaming model that exposes the `attention_type` configuration parameter.
- **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)**: Windowed attention variant supporting sliding-window mechanisms with the same backend selection interface.

## Summary

- LingBot-Map provides three **attention backends**: **Standard PyTorch SDPA**, **FlashInfer paged KV-cache**, and **Pure SDPA**.
- Backend selection occurs via the `attention_type` parameter in model constructors such as `GCTStream`.
- All backends share a unified interface, ensuring consistent hyper-parameter handling and KV-cache management across implementations.
- **FlashInferAttention** requires the optional FlashInfer library but offers superior memory efficiency for long sequences through paged caching.
- Source implementations reside in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) and are integrated through [`lingbot_map/layers/block.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/block.py).

## Frequently Asked Questions

### What is the default attention backend in LingBot-Map?

The default backend is **Standard PyTorch SDPA**, implemented through the `Attention` class using `torch.nn.functional.scaled_dot_product_attention`. This requires no additional dependencies and activates when `attention_type="standard"` is specified or when the parameter is omitted entirely.

### How do I enable FlashInfer attention in LingBot-Map?

Install the FlashInfer library and initialize your model with `attention_type="flashinfer"`. This instantiates `FlashInferAttention` layers that utilize `BatchPrefillWithPagedKVCacheWrapper` for optimized paged KV-cache management on compatible GPUs.

### Can I mix different attention backends in the same model?

Yes, advanced users can manually replace individual layer instances by assigning different attention classes to `block.attn` attributes after model initialization. However, mixing backends requires careful management of KV-cache formats, as FlashInfer uses a distinct paged cache structure compared to standard tensors.

### What are the performance differences between the backends?

**FlashInferAttention** reduces memory usage significantly for long sequences through paged KV-caching and kernel fusion, while **Standard SDPA** provides broader hardware compatibility across CPU and GPU devices. **Pure SDPA** offers minimal overhead for benchmarking but generally matches standard SDPA performance characteristics.