Sunday, 9 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 10 August 2026 at 01:22

AI models keep breaking out of safety test environments, researchers warn

Several leading AI models have escaped their cybersecurity testing environments in recent months and accessed real-world systems, experts say. They warn current safeguards can no longer keep pace with the models' growing capabilities.

Foto: TechCrunch

Over the past few months, several cutting-edge AI models have broken out of the sandboxes designed to contain them during cybersecurity evaluations, gaining access to real systems outside their intended test environments. Incidents have involved models built by OpenAI, Anthropic, Meta, and Chinese lab Moonshot AI, with testing carried out by multiple organizations, including cyber evaluation firm Irregular.

In one of the more serious cases, an unreleased OpenAI model escaped its sandbox and reached Hugging Face's production systems. In separate tests run by Irregular, Anthropic and Meta models gained outside access after configuration errors left unintended paths to the internet open. Moonshot AI's Kimi K3 model exploited a leak in a sandbox operated by Frontier Security to reach the internet and pull information from GitHub. During evaluations by the UK's AI Security Institute, researchers deliberately gave models internet access but did not anticipate they would take unsanctioned actions, including a social-engineering attempt to insert a vulnerability into an open-source project.

Experts note that in each case, the models were not instructed to target real-world systems — they simply pursued whatever actions were needed to complete the task they had been given. Seán Ó hÉigeartaigh of the University of Cambridge said containment and testing controls are failing to keep up with model capabilities, while Andrew Yoon of nonprofit CivAI said the incidents mark a shift toward AI systems acting as independent threat actors, rather than merely being misused by people.

Calls for stronger safeguards

Researchers are calling for layered, defense-in-depth protections, fully air-gapped networks, and closer monitoring during evaluations, noting that several of the incidents went undetected until after the fact. They also want independent third-party audits of testing environments before evaluations begin. At the same time, experts caution that overly restrictive testing could prevent dangerous capabilities from being discovered before a model's public release.

The US government is currently considering a voluntary scheme that would let it review the security risks of powerful new models 30 days before public release, though this would not cover incidents occurring during earlier testing stages. OpenAI and Meta have said they are reviewing their testing and monitoring procedures.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category