LingBot-Map Attention Backends: SDPA, FlashInfer, and Pure PyTorch Support
LingBot-Map supports three interchangeable attention backends—Standard PyTorch SDPA, FlashInfer paged KV-cache attention, and pure SDPA—allowing users to switch between them via the attention_type configuration parameter.
LingBot-Map is an open-source streaming transformer implementation that provides flexible attention mechanisms for efficient inference. The repository offers multiple attention backends to accommodate different hardware configurations and performance requirements, from standard PyTorch implementations to optimized FlashInfer kernels. Understanding these options enables developers to optimize memory usage and throughput for their specific deployment scenarios.
Supported Attention Backends in LingBot-Map
Standard PyTorch SDPA
The default backend utilizes PyTorch's native scaled_dot_product_attention through the Attention and CausalAttention classes defined in lingbot_map/layers/attention.py. This implementation requires no additional dependencies beyond standard PyTorch and supports both full-frame self-attention and causal streaming with KV-cache management. It serves as the fallback option when specialized hardware acceleration libraries are unavailable.
FlashInfer Paged KV-Cache Attention
For GPU-accelerated inference, LingBot-Map provides FlashInferAttention, which wraps FlashInfer's BatchPrefillWithPagedKVCacheWrapper to implement memory-efficient paged attention. This backend requires the optional FlashInfer library and is optimized for FP16/BF16 operations on compatible GPUs. The paged KV-cache approach significantly reduces memory fragmentation during long-sequence streaming inference, as implemented in lingbot_map/layers/flashinfer_cache.py.
Pure SDPA Attention
The SDPAAttention class offers a thin wrapper around PyTorch's SDPA that maintains API consistency with the other backends while isolating the attention computation. This design facilitates benchmarking comparisons and simplifies the integration of custom attention kernels, providing a clean interface that matches the hyper-parameter signatures of Attention and FlashInferAttention.
Configuring Attention Backends
The model architecture selects backends through the attention_type parameter in high-level model classes like GCTStream and GCTStreamWindow. All backend classes inherit from a shared abstract interface defined in the layers module, ensuring consistent handling of hyper-parameters including number of heads, head dimension, and dropout rates.
Default Standard Attention
from lingbot_map.models.gct_stream import GCTStream
# Use standard PyTorch SDPA (default)
model = GCTStream(attention_type="standard")
Enabling FlashInfer
from lingbot_map.models.gct_stream import GCTStream
# Enable FlashInfer paged attention (requires flashinfer package)
model = GCTStream(attention_type="flashinfer")
Pure SDPA Backend
from lingbot_map.models.gct_stream import GCTStream
# Use isolated SDPA wrapper
model = GCTStream(attention_type="sdpa")
Manual Backend Swapping
Advanced users can instantiate attention layers directly and replace them in existing blocks. This approach is useful for fine-grained control over specific transformer layers.
from lingbot_map.layers.attention import FlashInferAttention
from lingbot_map.layers.block import Block
# Create a transformer block with default attention
block = Block(dim=512, num_heads=8)
# Replace with FlashInfer attention manually
block.attn = FlashInferAttention(
dim=512,
num_heads=8,
fused_attn=True,
enable_3d_rope=False
)
Key Source Files
The attention backend implementations are organized across several core files:
lingbot_map/layers/attention.py: Contains theAttention,CausalAttention,FlashInferAttention, andSDPAAttentionclass definitions.lingbot_map/layers/flashinfer_cache.py: Provides utilities for managing FlashInfer's paged KV-cache buffers.lingbot_map/layers/block.py: Implements the transformerBlockclass that wires the selected attention backend into the architecture.lingbot_map/models/gct_stream.py: High-level streaming model that exposes theattention_typeconfiguration parameter.lingbot_map/models/gct_stream_window.py: Windowed attention variant supporting sliding-window mechanisms with the same backend selection interface.
Summary
- LingBot-Map provides three attention backends: Standard PyTorch SDPA, FlashInfer paged KV-cache, and Pure SDPA.
- Backend selection occurs via the
attention_typeparameter in model constructors such asGCTStream. - All backends share a unified interface, ensuring consistent hyper-parameter handling and KV-cache management across implementations.
- FlashInferAttention requires the optional FlashInfer library but offers superior memory efficiency for long sequences through paged caching.
- Source implementations reside in
lingbot_map/layers/attention.pyand are integrated throughlingbot_map/layers/block.py.
Frequently Asked Questions
What is the default attention backend in LingBot-Map?
The default backend is Standard PyTorch SDPA, implemented through the Attention class using torch.nn.functional.scaled_dot_product_attention. This requires no additional dependencies and activates when attention_type="standard" is specified or when the parameter is omitted entirely.
How do I enable FlashInfer attention in LingBot-Map?
Install the FlashInfer library and initialize your model with attention_type="flashinfer". This instantiates FlashInferAttention layers that utilize BatchPrefillWithPagedKVCacheWrapper for optimized paged KV-cache management on compatible GPUs.
Can I mix different attention backends in the same model?
Yes, advanced users can manually replace individual layer instances by assigning different attention classes to block.attn attributes after model initialization. However, mixing backends requires careful management of KV-cache formats, as FlashInfer uses a distinct paged cache structure compared to standard tensors.
What are the performance differences between the backends?
FlashInferAttention reduces memory usage significantly for long sequences through paged KV-caching and kernel fusion, while Standard SDPA provides broader hardware compatibility across CPU and GPU devices. Pure SDPA offers minimal overhead for benchmarking but generally matches standard SDPA performance characteristics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →