Kimi K3 Reinvents Transformer Residual Connections with Attention Residuals – First Change Since ResNet 2015
NewsSource: xAuthor: _avichawlaHotness: 126Published Jul 22, 2026
Attention Residuals replace fixed residual weights with softmax attention across layers, enabling input-dependent layer aggregation. Validated on Kimi's 48B model, Block AttnRes matches a 1.25x compute baseline with under 2% inference overhead. K3 scales to 2.8T parameters and pairs with Kimi Delta Attention for 6.3x faster decoding at million-token contexts. API live, weights open-source July 27.
- 开源模型
- Transformer
- Moonshot AI
- 深度学习
- Kimi K3
- 2.8T参数
- Attention Residuals
- 残差连接
Comments
Log in to comment
No comments yet. Be the first.