AI News
AI community posts, news and release updates
Latest updates
Anthropic releases Claude Opus 5, which surpasses its previous flagship Fable 5 on nearly all evaluation benchmarks, sparking excitement in the community.
- benchmark
- Fable 5
- Claude Opus 5
- Tutorial
Recreating Fall Guys with Same Prompt: Kimi K3 Beats Claude Opus 5.0 in Speed and Playability
Published Jul 24, 2026A developer used the same prompt to recreate Fall Guys with Kimi K3 and Claude Opus 5.0. Kimi K3 took ~9 min and $4.4, producing solid physics and gameplay. Opus 5.0 took ~17 min and $13+, with fancier UI but broken movement and lag. Kimi was far more playable.
- comparison
- game development
- Kimi K3
- Claude Opus 5.0
- Launch
Breaking: Anthropic Unveils Claude Opus 5, Near-Fable 5 Intelligence at Half the Price
Published Jul 24, 2026Anthropic officially launches Claude Opus 5, claiming it delivers intelligence close to Fable 5 at half the price, offering a more cost-effective option for developers.
- price
- Fable 5
- Claude Opus 5
- Idea
Open-Source and Frontier Labs Both Surge, but Opus Remains My Daily Driver
Published Jul 24, 2026The author acknowledges the rise of open-source models like Kimi and frontier progress from closed labs, supporting both. However, Claude Opus remains their daily driver for its thoughtfulness, capability, and unmatched style and personality, though they wish Anthropic invested more in frontend experience.
- open-source
- Claude Opus
- Kimi K3
- frontier progress
Artificial Analysis provides a comprehensive review of Claude Opus 5 across performance, pricing, and availability.
- analysis
- Claude Opus 5
- Launch
Opus 5 Sets New SOTA on ARC-AGI-3 at 30%, Demonstrating Strong Zero-Shot Problem Solving
Published Jul 24, 2026Claude Opus 5 achieves a new SOTA of 30% on ARC-AGI-3, a benchmark measuring problem-solving with no prior exposure—a setting where scaling typically yields the least improvement, making this jump impressive.
- SOTA
- Claude Opus 5
- ARC-AGI-3
- Launch
ChatGPT Work Agent Can Now Log Into Websites: Take Over Cloud Browser for Login
Published Jul 24, 2026OpenAI updates ChatGPT Work agent to handle websites requiring login; users can take over a cloud browser to sign in, with login persisting across sessions.
- ChatGPT
- 自动化
- Work agent
- 浏览器登录
Discusses technical details, effects, and risks of inserting or modifying system messages mid-conversation, valuable for complex dialogue applications.
- system messages
- conversation design
- Launch
CursorBench Results: Opus 5 Max Scores 70.0% at Half the Price of Fable 5 Max (70.5%)
Published Jul 24, 2026New CursorBench results show Opus 5 Max scoring 70.0%, just 0.5 points behind Fable 5 Max at 70.5%, but costing only $8.23 per task—less than half of Fable 5's $17.32. Moreover, Opus 5 Extra High beats Fable 5 Extra High outright, costing $4 less per task. Anthropic has made its own best model obsolete on value.
- Cost Efficiency
- Fable 5
- CursorBench
- Claude Opus 5
- Discussion
Jensen Huang and Scholar Agree: Open Models Must Be Tested on Capability, Control, Incentives, and Risk
Published Jul 24, 2026Jensen Huang aligns with @JennyQTa7 on the importance of open models and sovereign AI. Jenny's paper proposes a universal test for all models, including Kimi: capability, control, incentives, and risk.
- Kimi
- Jensen Huang
- AI policy
- open models