AI News
AI community posts, news and release updates
Latest updates
Baseten's Model APIs provide fast, reliable access to Kimi K3.
- API
- Baseten
- Kimi K3
Kimi K3 achieves up to 370 tokens per second on the vLLM inference framework, demonstrating high efficiency.
- vLLM
- Performance
- Kimi K3
Modal trained a custom DFlash speculator for Kimi K3, delivering faster inference without quality loss.
- inference acceleration
- Kimi K3
- Modal
- speculator
A former Anthropic employee alleges the company reduced AI safety safeguards in favor of performance, sparking debate about balancing safety and commerce.
- Anthropic
- AI safety
- safeguards
- whistleblower
The Kimi K3 technical report is out, detailing model architecture, training methods, and performance evaluation.
- Kimi K3
- technical report
The Kimi K3 team developed MiniTriton, a lightweight Triton-like compiler for efficient compilation of deep learning models.
- compiler
- Kimi K3
- miniTriton
- Launch
MoonEP open-sourced: high-performance communication library for distributed MoE
Published Jul 27, 2026MoonEP is designed to make expert-parallel communication more efficient at scale, reducing overhead in large MoE systems.
- open-source
- MoE
- MoonEP
- communication library
- Launch
AgentENV open-sourced: distributed system for running agent environments at scale
Published Jul 27, 2026AgentENV powers agentic RL training for Kimi K3 with fast snapshot, resume, and fork for large-scale parallel agent workflows.
- open-source
- RL
- agent environment
- AgentENV
- Launch
FlashKDA open-sourced: high-performance CUTLASS-based attention kernels with 1.72-2.22x prefill speedup
Published Jul 27, 2026FlashKDA is a CUTLASS-based implementation of Kimi Delta Attention, delivering 1.72-2.22x prefill speedup on H20 and works as a drop-in backend for flash-linear-attention.
- open-source
- FlashKDA
- attention kernel
- Kimi Delta Attention
The team launched an API for Kimi K3 on the release day, enabling immediate integration for developers.
- API
- Kimi K3
- day-0