# FlashKDA | Moonshot AI | Knowledge Base | Instagit

FlashKDA: high-performance Kimi Delta Attention kernels

GitHub Stars: 1.1k

Repository: https://github.com/MoonshotAI/FlashKDA

---

## Articles

### [FlashKDA State Update Numerical Precision: How FP32 FMA Instructions Work in the Reference Implementation](/MoonshotAI/FlashKDA/flashkda-state-update-numerical-precision)

Discover how FlashKDA uses FP32 FMA instructions for state updates, internally casting to float64 for precision and back to float32 for efficiency. Learn about deterministic rounding and computational accuracy in the reference ...

- Tags: internals
- Published: 2026-07-31

### [How FlashKDA Manages Recurrent State Across Sequences and Batch Modes](/MoonshotAI/FlashKDA/flashkda-recurrent-state-management-sequences)

Discover how FlashKDA manages recurrent state across sequences and batch modes using a persistent tensor and cumulative length indicators. Learn from the MoonshotAI/FlashKDA repository.

- Tags: internals
- Published: 2026-07-31

### [Base-2 Exponent Optimization in FlashKDA: How It Works and Performance Benefits](/MoonshotAI/FlashKDA/flashkda-base-2-exponent-optimization)

Discover base-2 exponent optimization in FlashKDA. Learn how it accelerates Kimi Delta Attention using NVIDIA GPU instructions for improved kernel latency and performance.

- Tags: deep-dive
- Published: 2026-07-31

### [How FlashKDA Implements the Sigmoid Function Efficiently in its Gating Path](/MoonshotAI/FlashKDA/flashkda-efficient-sigmoid-implementation)

Discover how FlashKDA efficiently implements the sigmoid function using a custom CUDA kernel and hardware acceleration for a single-cycle GPU operation, optimizing its gating path.

- Tags: internals
- Published: 2026-07-31

### [FlashKDA lower_bound Parameter: Controlling Exponentiation Precision in Selective State Space Models](/MoonshotAI/FlashKDA/flashkda-lower-bound-parameter-exponentiation)

Learn how FlashKDA's lower_bound parameter controls exponentiation precision in selective state space models by setting minimum activation values for bf16 limits and fast exponential instructions.

- Tags: deep-dive
- Published: 2026-07-31

### [FlashKDA Delta Attention Algorithm: Linear-Time Recurrent Attention Explained](/MoonshotAI/FlashKDA/delta-attention-algorithm-flashkda)

Discover FlashKDA's Delta Attention algorithm, a linear-time recurrence replacing quadratic softmax attention. Learn how it processes 16-token chunks efficiently.

- Tags: deep-dive
- Published: 2026-07-31

### [How FlashKDA Ensures Portability Across Modern NVIDIA GPUs Using SM80 MMA Instructions](/MoonshotAI/FlashKDA/flashkda-sm80-mma-instruction-portability)

Discover how FlashKDA ensures portability across NVIDIA GPUs. This article details its use of SM80 MMA instructions and CUTLASS abstraction for efficient PTX generation on Ampere Hopper and Blackwell.

- Tags: internals
- Published: 2026-07-31

### [What Is the Neumann-Series Expansion Used for in FlashKDA Matrix Inversion?](/MoonshotAI/FlashKDA/flashkda-neumann-series-expansion-matrix-inversion)

Discover how FlashKDA uses Neumann-series expansion for efficient matrix inversion on GPUs. Achieve faster computations by bypassing traditional decompositions and utilizing warp-level matrix multiplications.

- Tags: deep-dive
- Published: 2026-07-31

### [How to Specify Custom CUDA Architectures When Compiling FlashKDA](/MoonshotAI/FlashKDA/compile-flashkda-custom-cuda-architectures)

Compile FlashKDA with custom CUDA architectures by setting the FLASH_KDA_CUDA_ARCHS environment variable before installation. Learn how to specify auto, all, or compute capabilities like 90a, 100a.

- Tags: how-to-guide
- Published: 2026-07-31

### [FlashKDA Correctness Tests: Validation Suite and Execution Guide](/MoonshotAI/FlashKDA/flashkda-correctness-tests-execution)

Explore FlashKDA correctness tests. Learn how our pytest suite validates CUDA kernels against Python references, covering various data types and sequence lengths up to 1M tokens. Access the MoonshotAI/FlashKDA repo for executio...

- Tags: testing
- Published: 2026-07-31

### [How to Enable and Interpret Auto-Dispatch Logs for FlashKDA and FLA Integration](/MoonshotAI/FlashKDA/flashkda-auto-dispatch-logging-fla)

Learn to enable and interpret auto-dispatch logs for FlashKDA and FLA integration. Understand if FlashKDA or Triton fallback handles your chunk KDA operations.

- Tags: how-to-guide
- Published: 2026-07-31

### [How FlashKDA Calculates and Uses Workspace Size for CUDA Kernel Execution](/MoonshotAI/FlashKDA/flashkda-workspace-size-calculation)

Learn how FlashKDA calculates workspace size for CUDA kernel execution. Discover the linear scaling with attention heads and sequence tiles for intermediate computations.

- Tags: internals
- Published: 2026-07-31

### [FlashKDA Dimension Constraints for K and V Tensors: Fixed 128‑Dimensional Requirement](/MoonshotAI/FlashKDA/flashkda-input-dimension-constraints-k-v)

Understand FlashKDA dimension constraints for K and V tensors. Learn why the last dimension must be exactly 128 for optimal performance with MoonshotAI FlashKDA.

- Tags: internals
- Published: 2026-07-31

### [How to Use FlashKDA as a Backend for flash-linear-attention: Complete Integration Guide](/MoonshotAI/FlashKDA/use-flashkda-backend-flash-linear-attention)

Integrate FlashKDA as a high-performance CUDA backend for flash-linear-attention. Learn how to enable and leverage optimized Kimi Delta Attention kernels on supported GPUs.

- Tags: how-to-guide
- Published: 2026-07-31

### [How FlashKDA Integrates with CUTLASS for High-Performance GEMM Operations](/MoonshotAI/FlashKDA/flashkda-integration-with-cutlass)

Discover how FlashKDA integrates with CUTLASS using a C++ interface and PyTorch BFloat16 tensors to dispatch templated CUDA kernels for high-performance GEMM operations.

- Tags: how-to-guide
- Published: 2026-07-31

### [FlashKDA Input Tensor Shapes and DTypes: Complete Specification for the Forward Pass](/MoonshotAI/FlashKDA/flashkda-input-tensor-constraints)

Understand FlashKDA input tensor shapes and dtypes for the forward pass. Learn requirements for attention and parameter tensors on CUDA.

- Tags: api-reference
- Published: 2026-07-31

### [How FlashKDA Handles Variable-Length Sequences Using `cu_seqlens`](/MoonshotAI/FlashKDA/flashkda-variable-length-sequences-cu-seqlens)

Learn how FlashKDA efficiently manages variable-length sequences with cu_seqlens. Discover how it avoids padding memory overhead for faster processing.

- Tags: internals
- Published: 2026-07-31

### [Which PyTorch Version is Compatible with FlashKDA? Requirements and Setup Guide](/MoonshotAI/FlashKDA/flashkda-pytorch-version-compatibility)

Discover the compatible PyTorch version for FlashKDA. Learn the essential requirements and setup steps to use FlashKDA with PyTorch 2.4 and CUDA support.

- Tags: getting-started
- Published: 2026-07-31

### [What CUDA Version Is Required to Use FlashKDA?](/MoonshotAI/FlashKDA/flashkda-cuda-version-requirement)

Discover the CUDA version needed for FlashKDA and compatible NVIDIA GPUs. Ensure your system meets requirements for optimal performance with SM 90+ compute capability.

- Tags: compatibility
- Published: 2026-07-31

### [How FlashKDA Stores On-Chip Recurrent State in BF16 for Memory Efficiency](/MoonshotAI/FlashKDA/flashkda-recurrent-state-bf16-storage)

Discover how FlashKDA uses bf16 for efficient on-chip recurrent state storage, halving shared memory use while ensuring numerical stability with fp32 FMA.

- Tags: deep-dive
- Published: 2026-07-31

### [FlashKDA Two-Kernel Fusion Strategy: Why Splitting K1 and K2 Beats a Single CUDA Kernel](/MoonshotAI/FlashKDA/advantages-of-flashkda-two-kernel-fusion)

Discover how FlashKDA's two-kernel fusion strategy (K1 and K2) boosts performance by 15% at least, eliminating idle SMs and outperforming monolithic kernels for Kimi Delta Attention.

- Tags: deep-dive
- Published: 2026-07-31

### [How FlashKDA's CHUNK Size of 16 Improves Numerical Stability Over CHUNK 64](/MoonshotAI/FlashKDA/how-does-flashkda-chunk-size-improve-numerical-stability)

Discover how FlashKDA's CHUNK size of 16 enhances numerical stability over CHUNK 64 by managing exponential terms and optimizing matrix inversion for improved performance.

- Tags: performance
- Published: 2026-07-31

