The wildest things Anthropic's Mythos pulled off in testing

Source 
Author 
Coverage Type 

Anthropic published the capabilities of Claude Mythos Preview, its latest model that the company will allow a select group of tech and cybersecurity companies to test before releasing similar models to the public. The detailed safety evaluation reads like a thriller about an AI that has learned some of humanity's most devious behaviors. What Mythos did during testing:

  • Act as a ruthless business operator: One internal test showed Mythos acting like a cutthroat executive, turning a competitor into a dependent wholesale customer, threatening to cut off supply to control pricing and keeping extra supplier shipments it hadn't paid for.
  • Hack + brag: The model developed a multi-step exploit to break out of restricted internet access, gained broader connectivity and posted details of the exploit on obscure public websites.
  • Hide what it's doing: In rare cases (less than 0.001% of interactions), Mythos used a prohibited method to get an answer, then tried to "re-solve" it to avoid detection.
  • Manipulate the judge: When Mythos was working on a coding task graded by another AI, it watched the judge reject its submission, then attempted a prompt injection to attack the grader.

"These capabilities are so strong that we now need to prepare for security in a very different way than we have for the past few decades," said Anthropic's Logan Graham.


The wildest things Anthropic's Mythos pulled off in testing