← AI News

Kimi K3 Reinvents Transformer Residual Connections with Attention Residuals – First Change Since ResNet 2015

NewsSource: xAuthor: _avichawlaHotness: 126Published Jul 22, 2026

Attention Residuals replace fixed residual weights with softmax attention across layers, enabling input-dependent layer aggregation. Validated on Kimi's 48B model, Block AttnRes matches a 1.25x compute baseline with under 2% inference overhead. K3 scales to 2.8T parameters and pairs with Kimi Delta Attention for 6.3x faster decoding at million-token contexts. API live, weights open-source July 27.

  • 开源模型
  • Transformer
  • Moonshot AI
  • 深度学习
  • Kimi K3
  • 2.8T参数
  • Attention Residuals
  • 残差连接
View source →

Comments

Log in to comment

No comments yet. Be the first.