← AI News

FlashKDA open-sourced: high-performance CUTLASS-based attention kernels with 1.72-2.22x prefill speedup

LaunchSource: xAuthor: Kimi_MoonshotHotness: 2128Published Jul 27, 2026

FlashKDA is a CUTLASS-based implementation of Kimi Delta Attention, delivering 1.72-2.22x prefill speedup on H20 and works as a drop-in backend for flash-linear-attention.

  • open-source
  • FlashKDA
  • attention kernel
  • Kimi Delta Attention
View source →

Comments

Log in to comment

No comments yet. Be the first.