Analysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8%

Can you trust an AI model to do what you intended? This is a central question both for those deploying AI systems...