FlashKDA

FlashKDA: high-performance Kimi Delta Attention kernels

22 articles 1.1k View on GitHub ↗
22 articles
FlashKDA State Update Numerical Precision: How FP32 FMA Instructions Work in the Reference Implementation

Discover how FlashKDA uses FP32 FMA instructions for state updates, internally casting to float64 for precision and back to float32 for efficiency. Learn about deterministic rounding and computational accuracy in the reference ...

internals
Jul 31, 2026
How FlashKDA Manages Recurrent State Across Sequences and Batch Modes

Discover how FlashKDA manages recurrent state across sequences and batch modes using a persistent tensor and cumulative length indicators. Learn from the MoonshotAI/FlashKDA repository.

internals
Jul 31, 2026
Base-2 Exponent Optimization in FlashKDA: How It Works and Performance Benefits

Discover base-2 exponent optimization in FlashKDA. Learn how it accelerates Kimi Delta Attention using NVIDIA GPU instructions for improved kernel latency and performance.

deep-dive
Jul 31, 2026
How FlashKDA Implements the Sigmoid Function Efficiently in its Gating Path

Discover how FlashKDA efficiently implements the sigmoid function using a custom CUDA kernel and hardware acceleration for a single-cycle GPU operation, optimizing its gating path.

internals
Jul 31, 2026
FlashKDA lower_bound Parameter: Controlling Exponentiation Precision in Selective State Space Models

Learn how FlashKDA's lower_bound parameter controls exponentiation precision in selective state space models by setting minimum activation values for bf16 limits and fast exponential instructions.

deep-dive
Jul 31, 2026
FlashKDA Delta Attention Algorithm: Linear-Time Recurrent Attention Explained

Discover FlashKDA's Delta Attention algorithm, a linear-time recurrence replacing quadratic softmax attention. Learn how it processes 16-token chunks efficiently.

deep-dive
Jul 31, 2026
How FlashKDA Ensures Portability Across Modern NVIDIA GPUs Using SM80 MMA Instructions

Discover how FlashKDA ensures portability across NVIDIA GPUs. This article details its use of SM80 MMA instructions and CUTLASS abstraction for efficient PTX generation on Ampere Hopper and Blackwell.

internals
Jul 31, 2026
What Is the Neumann-Series Expansion Used for in FlashKDA Matrix Inversion?

Discover how FlashKDA uses Neumann-series expansion for efficient matrix inversion on GPUs. Achieve faster computations by bypassing traditional decompositions and utilizing warp-level matrix multiplications.

deep-dive
Jul 31, 2026
How to Specify Custom CUDA Architectures When Compiling FlashKDA

Compile FlashKDA with custom CUDA architectures by setting the FLASH_KDA_CUDA_ARCHS environment variable before installation. Learn how to specify auto, all, or compute capabilities like 90a, 100a.

how-to-guide
Jul 31, 2026
FlashKDA Correctness Tests: Validation Suite and Execution Guide

Explore FlashKDA correctness tests. Learn how our pytest suite validates CUDA kernels against Python references, covering various data types and sequence lengths up to 1M tokens. Access the MoonshotAI/FlashKDA repo for executio...

testing
Jul 31, 2026
How to Enable and Interpret Auto-Dispatch Logs for FlashKDA and FLA Integration

Learn to enable and interpret auto-dispatch logs for FlashKDA and FLA integration. Understand if FlashKDA or Triton fallback handles your chunk KDA operations.

how-to-guide
Jul 31, 2026
How FlashKDA Calculates and Uses Workspace Size for CUDA Kernel Execution

Learn how FlashKDA calculates workspace size for CUDA kernel execution. Discover the linear scaling with attention heads and sequence tiles for intermediate computations.

internals
Jul 31, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →