AI News
AI community posts, news and release updates
Latest updates
- Launch
Opus 5 Sets New SOTA on ARC-AGI-3 at 30%, Demonstrating Strong Zero-Shot Problem Solving
Published Jul 24, 2026Claude Opus 5 achieves a new SOTA of 30% on ARC-AGI-3, a benchmark measuring problem-solving with no prior exposure—a setting where scaling typically yields the least improvement, making this jump impressive.
- SOTA
- Claude Opus 5
- ARC-AGI-3
- Launch
ChatGPT Work Agent Can Now Log Into Websites: Take Over Cloud Browser for Login
Published Jul 24, 2026OpenAI updates ChatGPT Work agent to handle websites requiring login; users can take over a cloud browser to sign in, with login persisting across sessions.
- ChatGPT
- 自动化
- Work agent
- 浏览器登录
Discusses technical details, effects, and risks of inserting or modifying system messages mid-conversation, valuable for complex dialogue applications.
- system messages
- conversation design
- Launch
CursorBench Results: Opus 5 Max Scores 70.0% at Half the Price of Fable 5 Max (70.5%)
Published Jul 24, 2026New CursorBench results show Opus 5 Max scoring 70.0%, just 0.5 points behind Fable 5 Max at 70.5%, but costing only $8.23 per task—less than half of Fable 5's $17.32. Moreover, Opus 5 Extra High beats Fable 5 Extra High outright, costing $4 less per task. Anthropic has made its own best model obsolete on value.
- Cost Efficiency
- Fable 5
- CursorBench
- Claude Opus 5
- Discussion
Jensen Huang and Scholar Agree: Open Models Must Be Tested on Capability, Control, Incentives, and Risk
Published Jul 24, 2026Jensen Huang aligns with @JennyQTa7 on the importance of open models and sovereign AI. Jenny's paper proposes a universal test for all models, including Kimi: capability, control, incentives, and risk.
- Kimi
- Jensen Huang
- AI policy
- open models
The author admits they were wrong: Opus 5 outperforms Fable 5 on nearly every benchmark at half the cost, recommending it as the default in Claude Code.
- Claude Code
- 对比
- Fable 5
- Opus 5
- 成本效率
Claude Opus 5 has entered Agent Arena, where it will be evaluated on millions of real-world agentic tasks using web search, filesystem, and terminal tools. Scores coming soon.
- Agent Arena
- 性能测试
- Claude Opus 5
Anthropic has just announced the launch of the new Claude Opus 5 AI model.
- 发布
- Claude Opus 5
Anthropic announces major updates for Claude Opus 5, including longer context windows, improved reasoning, and new safety features.
- New Features
- Claude Opus 5
- Rant
Opus 5 beats Fable 5 in every benchmark but didn't need government approval? 😭
Published Jul 24, 2026A user complains that Opus 5 outperforms Fable 5 in every benchmark yet avoided the extensive regulatory approvals that Fable 5 required.
- 吐槽
- Fable 5
- Opus 5