← AI News

Composite-Bench released: long-horizon computer-use benchmark, GLM-5.2 beats Kimi K3 by 32 points, surpasses all closed models except Claude

LaunchSource: xAuthor: YangFanYunHotness: 359Published Jul 24, 2026

Today sees the release of Composite-Bench, a long-horizon computer-use benchmark with certified-optimal answers and verified compute. The strongest open-weights model is not Kimi K3; GLM-5.2 beats it by 32 points and outperforms every closed model tested except Claude.

  • benchmark
  • GLM 5.2
  • Computer Use
  • Kimi K3
  • Composite-Bench
View source →

Comments

Log in to comment

No comments yet. Be the first.