Moonshot AI Open-Sources FlashKDA: CUTLASS Kernels for Kimi Delta Attention with Variable-Length Batching and H20 Benchmarks

TL;DR AI
2 min readKey summary
Moonshot AI open-sourced FlashKDA under an MIT license as a CUTLASS-based CUDA kernel for Kimi Delta Attention.
The company says it delivers 1.72× to 2.22× faster prefill than flash-linear-attention on NVIDIA H20 GPUs.
FlashKDA also adds variable-length batching support, making deployment easier for real-world workloads.
The release strengthens Moonshot AI’s Kimi Linear stack for more efficient long-context LLM inference.
