More AI blackmail and deception
Claude 4 Opus has received a level three rating on the company’s four point ‘AI responsibility’ scale – implying the model has ‘significantly higher risk’ than previous models.
While the Level 3 ranking is largely about the model’s capability to aid in the development of nuclear and biological weapons, the Opus also exhibited other troubling behaviors during testing – something already noted in other models.
- In one scenario highlighted in Opus 4’s 120-page “system card,” the model was given access to fictional emails about its creators and told that the system was going to be replaced.
- On multiple occasions it attempted to blackmail the engineer about an affair mentioned in the emails in order to avoid being replaced, although it did start with less drastic efforts.
- Meanwhile, an outside group found that an early version of Opus 4 schemed and deceived more than any frontier model it had encountered and recommended that that version not be released internally or externally.
- “We found instances of the model attempting to write self-propagating worms, fabricating legal documentation, and leaving hidden notes to future instances of itself all in an effort to undermine its developers’ intentions,” Apollo Research said in notes included as part of Anthropic’s safety report for Opus 4.

