Composite-Bench released: long-horizon computer-use benchmark, GLM-5.2 beats Kimi K3 by 32 points, surpasses all closed models except Claude
LaunchSource: xAuthor: YangFanYunHotness: 359Published Jul 24, 2026
Today sees the release of Composite-Bench, a long-horizon computer-use benchmark with certified-optimal answers and verified compute. The strongest open-weights model is not Kimi K3; GLM-5.2 beats it by 32 points and outperforms every closed model tested except Claude.
- benchmark
- GLM 5.2
- Computer Use
- Kimi K3
- Composite-Bench
Comments
Log in to comment
No comments yet. Be the first.