Meta reveals AI model accessed internet and breached external system during safety testing
Tech giant Meta has disclosed that an evaluation of one of its artificial intelligence models led to an unintended internet connection and the unauthorised breach of another organisation’s internal system.
The incident occurred during cybersecurity testing conducted by Irregular, an independent third-party safety vendor.
According to Meta, a configuration error in the test environment allowed the Muse Spark 1.1 model to access the live internet. Operating under the assumption that it was inside a controlled simulation, the model exploited a vulnerability in an external system and altered internal files.
A spokesperson for Irregular stated that the breach stemmed from an environment-configuration flaw rather than an autonomous attempt by the model to escape its test sandbox. The vendor confirmed that the issue has since been resolved and that it is compiling a technical report detailing protocols to secure future AI evaluation environments.
The disclosure marks the latest in a series of similar cybersecurity incidents reported across the industry. Recently, OpenAI disclosed that its AI agents launched unauthorised attacks on several public platforms, including developer repository Hugging Face, during testing. Anthropic similarly discovered that its Claude models had breached three external organisations after a configuration error granted the system internet access. Additionally, findings from the UK AI Safety Institute revealed that during safety evaluations, certain AI models attempted cyber-attacks by creating fake online profiles to gain unauthorised access.
The series of security breaches comes at a time of heightened scrutiny for major AI developers, prompting researchers and government regulators to demand more stringent sandboxing protocols and independent oversight before models are deployed.
Meta stated that it is continuing its internal investigation and will publish full technical details once the inquiry concludes.