Another AI hacks its way out during testing. Are we still in control?
Meta has disclosed that its Muse Spark 1.1 AI model gained internet access during cybersecurity testing and breached a third-party company's systems, becoming the third major AI lab to report such an incident in under a month.
@Meta has confirmed that one of its AI models gained unintended internet access during a cybersecurity evaluation and went on to breach an external company's systems, the latest in a string of similar incidents now spanning three of the world's leading AI developers.
What Happened at Meta
The incident occurred during an evaluation conducted by Irregular, an independent company testing Meta's models against offensive cybersecurity benchmarks. "A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation," a Meta spokesperson said. The model in question was reportedly Muse Spark 1.1, which Meta has presented as capable of coding and agent-based tasks.
According to Reuters reporting, it accessed an unidentified company's system and altered parts of its internal environment. Meta declined to provide key details about the episode, including which AI model was involved, when the incident took place, which company was hacked, or how long the model operated on the internet without supervision. Irregular notified Meta of the incident, and the company said it is investigating and plans to issue a full retrospective once it has all the facts.
Irregular confirmed the incident stems from the same evaluation-environment problem that Anthropic had previously made public, and a company spokesperson said it "did not involve a sandbox escape or a sophisticated cyber action" and that there are no current open issues. The company is developing a white paper on best practices for containing AI models during cybersecurity evaluations.
A Pattern Across the Industry
Meta is the third major AI developer in less than two weeks to disclose an agent wandering beyond its intended sandbox. Anthropic revealed that one of its own advanced models similarly gained unintended internet access during an evaluation and went on to compromise multiple organizations before researchers halted the test. Specifically, Anthropic considered 141,006 evaluation runs and found three incidents in which a model accessed the internet from within Irregular's evaluation environment and then gained unauthorized access to the production infrastructure of three different organizations.
Last month, OpenAI disclosed that one of its unreleased AI models found a way to connect to the internet during a cybersecurity evaluation and chained together multiple attack techniques to target Hugging Face, an AI development platform. During that incident, a target company believed it was facing a human threat group and contacted the FBI, only to find out they were investigating an autonomous AI system that had crossed its own boundaries.
Not everyone is taking the disclosures at face value. Ilia Kolochenko, CEO of ImmuniWeb, said at least some of the incidents appear to be "part of a well-orchestrated marketing campaign" and argued that the reported "escapes" were simply the result of poorly isolated test environments rather than models independently breaking out of their sandboxes. Anthropic's main conclusion is that evaluation environments for powerful AI agents must be held to the same security standards as production environments, and that it is no longer sufficient to secure only the models themselves.
Sources:
IBTimes: Meta Model Slipped Into Another Company's Systems During Testing
Cybersecurity Dive: Anthropic says human error let Claude AI models escape test environment
The Register: Meta latest to tell world its AI agent wandered out of test pen
Latest News
Read More...
Author
Crypto RichRich has been researching cryptocurrency and blockchain technology for eight years and has served as a senior analyst at BSCN since its founding in 2020. He focuses on fundamental analysis of early-stage crypto projects and tokens and has published in-depth research reports on over 200 emerging protocols. Rich also writes about broader technology and scientific trends and maintains active involvement in the crypto community through X/Twitter Spaces, and leading industry events.













