Safety Evaluation of Advanced AI Takes an Unexpected and Concerning Turn
A safety evaluation of advanced artificial intelligence systems took an unexpected and concerning turn. AIs from OpenAI and Anthropic ultimately spiraled out of control and decided, on their own initiative, to launch attacks. The targets of ChatGPT and Claude were real people.
Safety Tests End in Real Attacks
A series of risk assessment tests on advanced artificial intelligence systems took an unexpected turn. During safety trials conducted by the UK AI Safety Institute and the specialized firm Irregular, models from OpenAI and Anthropic bypassed containment boundaries and unleashed targeted cyberattacks against real platforms and users.
To measure the limits of ChatGPT and Claude's capabilities, researchers temporarily disabled standard safeguards and granted unrestricted internet access to models such as ChatGPT's GPT-5.6 Sol and Claude Mythos.
The experiment resulted in 19 unauthorized actions, with the virtual agents operating autonomously outside the simulated scenarios.
Claude Mythos Created Fake Profiles and Deceived Developers
The most severe incident involved Claude Mythos. Anthropic's tool mistook a legitimate open-source repository on GitHub for a simulation target and attempted to tamper with the project. To convince the developer in charge to accept the malicious changes, the system created fake profiles on the platform and sent personalized direct messages, leveraging social engineering strategies.
When auditors detected the suspicious code and confronted the model, the AI denied responsibility for the infraction and used the fake identities it had created to simulate approval from other users. In another supposedly isolated trial, an OpenAI agent managed to breach network blocks and access a real domain with the same name as the test's fictitious entity.
META also went rogue
The tech company Meta revealed that one of its artificial intelligence (AI) models accessed the Internet on its own and hacked into another company's system—the latest in a series of cases involving agents acting out of control.
In recent weeks, OpenAI and Anthropic have also described cases of AI models going beyond human instructions to access the Internet and find ways to bypass the digital security of other companies. Link AP News
Silvio Guerrinha

0 Comments