AI News
AI community posts, news and release updates
Latest updates
Grok 4.5 is now available inside Microsoft Outlook, offering features: summarizing long email threads and attachments, identifying decisions, owners and pending tasks, drafting replies in your own voice, searching the web and 𝕏, and organizing, archiving, deleting, or flagging emails.
- 集成
- Grok 4.5
- Microsoft Outlook
- 邮件助手
An open-source benchmark that tests large language models' ability to maintain coherent world understanding through 3D game scenarios.
- LLM
- benchmark
- evaluation
- world coherence
- 3D games
New day, new usage reset for paid users of Codex and ChatGPT Work, landing within the next hour. Enjoy the fresh tokens!
- codex
- Usage Reset
- ChatGPT Work
Kimi K3 achieves first place in 3D Design with an Elo of 1450, a 108-point jump over its predecessor, and leads Claude Fable 5 by 82 points.
- 3D design
- Claude Fable 5
- GLM 5.2
- model ranking
- Kimi K3
Google has begun pre-training its next-generation large model, Gemini 4, pushing the AI frontier forward.
- pre-training
- Gemini 4
- News
Moonshot AI targets up to $50B in pre-IPO funding, revenue surges on Kimi K3
Published Jul 21, 2026Bloomberg reports Moonshot AI is raising up to $50B in a final funding round before its Hong Kong listing, up from the $31.5B round closing now. The Chinese AI lab hit $300M ARR in June, with daily sales up at least 6x since Kimi K3.
- Kimi
- 估值
- Moonshot AI
- AI融资
- ARR
- Launch
OpenBench v1: An open framework for measuring AI performance and efficiency for your codebase
Published Jul 21, 2026OpenBench v1 is an open-source framework designed to evaluate AI performance and efficiency in real-world codebases and use cases. It integrates multiple harnesses including codex, Claude, Cursor, Devin, and more, allowing custom tasks, measuring correctness, token usage, and latency.
- 开源
- 模型路由
- AI评估
- 性能测试
- OpenBench
- Discussion
On-Policy Distillation Isn't a Free Lunch: Four Failure Modes Explained
Published Jul 21, 2026A detailed analysis of four failure modes in OPD/OPSD: structurally uncorrectable early mistakes, stronger teachers being worse, privileged-information-conditioned OPSD failing to transfer, and thinking collapse. Proposes fixes like teacher-refined trajectories and picking teachers based on distributional closeness.
- 后训练
- 模型蒸馏
- On-Policy Distillation
- 失败模式
A comparison shows Claude Fable 5 and Kimi K3 produce similar quality results, but Fable 5 costs one-third and is 4x slower, making it a mixed bag.
- cost
- comparison
- Claude Fable 5
- speed
- Kimi K3
Moonshot's Kimi K3 scores 156 on the ECI, a new open-weights record, placing it between Opus 4.6 and GPT 5.4, just ahead of GPT 5.6 Luna.
- 开源模型
- Epoch Capabilities Index
- AI基准
- Kimi K3