FlashKDA open-sourced: high-performance CUTLASS-based attention kernels with 1.72-2.22x prefill speedup
LaunchSource: xAuthor: Kimi_MoonshotHotness: 2128Published Jul 27, 2026
FlashKDA is a CUTLASS-based implementation of Kimi Delta Attention, delivering 1.72-2.22x prefill speedup on H20 and works as a drop-in backend for flash-linear-attention.
- open-source
- FlashKDA
- attention kernel
- Kimi Delta Attention
Comments
Log in to comment
No comments yet. Be the first.