AI News
AI community posts, news and release updates
Latest updates
- News
Moonshot AI limits Kimi K3 subscriptions due to compute constraints: inference bottleneck for 2.8T param model
Published Jul 21, 2026Moonshot AI is capping new Kimi K3 subscriptions because of compute limitations. The 2.8 trillion parameter model is trained, but scaling inference requires massive accelerators, memory, and networking. Bullish for MU, SKHY, AVGO, CRDO, ALAB, MRVL.
- memory
- networking
- AI hardware
- Kimi K3
- compute constraints
- inference bottleneck
A deep dive into distillation's role in post-training—SFT, RL, and why the best teacher isn't always the strongest model. Argues distillation is becoming less impactful over time.
- RL
- 蒸馏
- 模型训练
- 后训练
- SFT
- Launch
Usage limits doubled for individual and team plans: applies to Grok, Composer, and new Cursor models
Published Jul 21, 2026Cursor official announcement: usage limits doubled for all individual and team plans, covering Grok, Composer, and all new models.
- Cursor
- Grok
- Composer
- 使用限制翻倍
- Launch
Gemini 3.6 Flash is here: better at production code, avoids loops, excels at multimodal tasks
Published Jul 21, 2026Google launches Gemini 3.6 Flash, which writes production-ready code faster without getting stuck in loops, and excels at multimodal tasks like chart analysis, document understanding, and report drafting. Available in GeminiApp, Antigravity, Google AI Studio, and Android Studio.
- Gemini
- multimodal
- model release
- 3.6 Flash
- Launch
Gemini 3.6 Flash vs 3.5 Flash: higher quality, fewer tokens — watch the comparison
Published Jul 21, 2026Google shares a comparison video of Gemini 3.6 Flash against 3.5 Flash, showing improved quality and lower token usage at the same cost.
- Gemini
- comparison
- Quality
- 3.6 Flash
- Launch
Poolside Releases Laguna S 2.1: 118B-A8B Open-Weight Coding Agent Model with 1M Context
Published Jul 21, 2026Laguna S 2.1 achieves 70.2% on Terminal-Bench 2.1, beating larger models like Thinking Machines Inkling (975B) and DeepSeek V4 Pro Max (1.6T). It can build a complete HTML/CSS rendering engine in 50 minutes without human intervention.
- 开源模型
- 编码模型
- Laguna S 2.1
- 代码智能体
The SYNAPS-I project led by Berkeley Lab uses SAM 3 and DINOv3 to automate image segmentation, reducing 3D volume labeling from a month of manual effort to ~15 minutes, accelerating scientific discovery for the DOE's Genesis Mission.
- scientific discovery
- SAM 3
- DINOv3
- image segmentation
- Berkeley Lab
- News
Memory demand growing: Nvidia asks Samsung to boost V-NAND, Kimi K3 rejects new users due to capacity
Published Jul 21, 2026The author argues memory demand is still growing: Nvidia asked Samsung to scale up V9/V10 V-NAND; Chinese Kimi K3 had to turn away new subscribers due to capacity; TSMC report shows strong AI chip demand. Stocks are up but fundamentals are solid, unlike the dot-com bubble.
- Nvidia
- memory
- Samsung
- TSMC
- Kimi K3
- AI demand
- News
The Download: Chinese AI divides the White House, and a record copyright payout
Published Jul 21, 2026This issue of The Download covers how Chinese AI is causing divisions within the White House, along with a record-breaking copyright payout.
- copyright
- Chinese AI
- White House
- payout
A user bought the $20/month Qwen 3.8 Max plan, ran dozens of sub-agents, and used only 7% of the quota. The same workload would have maxed out Claude. The model is genuinely good—how can Anthropic compete?
- Claude
- Qwen
- 吐槽
- 性价比