Anthropic's Frontier Red Team evaluated the Chinese model GLM-5.3 on 100 tasks from an internal Binary Exploitation benchmark and found it capable of developing full control flow hijacks in 4% of trials.

  • GLM-5.3 succeeded in 4% of trials, while Claude Mythos Preview achieved this in 6%.
  • Earlier models, including Claude Opus 4.6 and GLM-5.2, did not succeed in any of the tasks.
  • The results indicate that a meaningful threshold has been crossed regarding advanced cyber capabilities.