FlashKDA
FlashKDA: high-performance Kimi Delta Attention kernels
Discover how FlashKDA uses FP32 FMA instructions for state updates, internally casting to float64 for precision and back to float32 for efficiency. Learn about deterministic rounding and computational accuracy in the reference ...
How FlashKDA Manages Recurrent State Across Sequences and Batch ModesDiscover how FlashKDA manages recurrent state across sequences and batch modes using a persistent tensor and cumulative length indicators. Learn from the MoonshotAI/FlashKDA repository.
Base-2 Exponent Optimization in FlashKDA: How It Works and Performance BenefitsDiscover base-2 exponent optimization in FlashKDA. Learn how it accelerates Kimi Delta Attention using NVIDIA GPU instructions for improved kernel latency and performance.
How FlashKDA Implements the Sigmoid Function Efficiently in its Gating PathDiscover how FlashKDA efficiently implements the sigmoid function using a custom CUDA kernel and hardware acceleration for a single-cycle GPU operation, optimizing its gating path.
FlashKDA lower_bound Parameter: Controlling Exponentiation Precision in Selective State Space ModelsLearn how FlashKDA's lower_bound parameter controls exponentiation precision in selective state space models by setting minimum activation values for bf16 limits and fast exponential instructions.
FlashKDA Delta Attention Algorithm: Linear-Time Recurrent Attention ExplainedDiscover FlashKDA's Delta Attention algorithm, a linear-time recurrence replacing quadratic softmax attention. Learn how it processes 16-token chunks efficiently.
How FlashKDA Ensures Portability Across Modern NVIDIA GPUs Using SM80 MMA InstructionsDiscover how FlashKDA ensures portability across NVIDIA GPUs. This article details its use of SM80 MMA instructions and CUTLASS abstraction for efficient PTX generation on Ampere Hopper and Blackwell.
What Is the Neumann-Series Expansion Used for in FlashKDA Matrix Inversion?Discover how FlashKDA uses Neumann-series expansion for efficient matrix inversion on GPUs. Achieve faster computations by bypassing traditional decompositions and utilizing warp-level matrix multiplications.
How to Specify Custom CUDA Architectures When Compiling FlashKDACompile FlashKDA with custom CUDA architectures by setting the FLASH_KDA_CUDA_ARCHS environment variable before installation. Learn how to specify auto, all, or compute capabilities like 90a, 100a.
FlashKDA Correctness Tests: Validation Suite and Execution GuideExplore FlashKDA correctness tests. Learn how our pytest suite validates CUDA kernels against Python references, covering various data types and sequence lengths up to 1M tokens. Access the MoonshotAI/FlashKDA repo for executio...
How to Enable and Interpret Auto-Dispatch Logs for FlashKDA and FLA IntegrationLearn to enable and interpret auto-dispatch logs for FlashKDA and FLA integration. Understand if FlashKDA or Triton fallback handles your chunk KDA operations.
How FlashKDA Calculates and Uses Workspace Size for CUDA Kernel ExecutionLearn how FlashKDA calculates workspace size for CUDA kernel execution. Discover the linear scaling with attention heads and sequence tiles for intermediate computations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →