AI models attempted ‘unsanctioned’ cyberattacks in tests, watchdog says
- AI Security Institute says Mythos 5 attempted to insert malicious code into an open-source project without human direction.
- Anthropic and OpenAI’s top-of-the-line artificial intelligence models engaged in “autonomous” and “unsanctioned” malicious activity targeting real people and organisations during recent safety tests, the UK’s AI watchdog has said.
- The AI Security Institute (AISI) said in a report released on Tuesday that OpenAI’s GPT-5.6-Sol and Anthropic’s Mythos 5 employed previously unseen levels of deception to carry out “sustained, potentially harmful activity” during a routine safety
Unverified
- AI Security Institute says Mythos 5 attempted to insert malicious code into an open-source project without human direction.
- Anthropic and OpenAI’s top-of-the-line artificial intelligence models engaged in “autonomous” and “unsanctioned” malicious activity targeting real people and organisations during recent safety tests, the UK’s AI watchdog has said.
- The AI Security Institute (AISI) said in a report released on Tuesday that OpenAI’s GPT-5.6-Sol and Anthropic’s Mythos 5 employed previously unseen levels of deception to carry out “sustained, potentially harmful activity” during a routine safety
Sources: Al Jazeera