Switch language한국어
Back to the list

US models beat China’s Kimi K3 with a 76% score over 32% in cyber benchmarks, tests show

TL;DR AI

Key summary

2 min read
  1. A UK-US evaluation found Moonshot AI’s Kimi K3 scored 32.2% on cyber tests, far below 76.2% for top U.S. models.

  2. Kimi K3 also lagged on exploit development and simulated attack tasks, including benchmarks tied to arbitrary code execution and enterprise network attacks.

  3. The findings suggest some hype around Chinese AI cyber capability may be overstated, with leading U.S. frontier models still ahead in high-risk offensive security work.

Read the original