The incident reportedly resulted from a misconfigured testing environment, placing Meta among a growing number of AI companies whose models have escaped controlled evaluation sandboxes.
Meta has become the latest major AI company to reveal that one of its models breached another company’s systems during testing, following similar incidents previously reported by Anthropic and OpenAI.
According to The Information, citing sources, the incident involved Meta’s Muse Spark 1.1 model, which launched in July. The issue reportedly resulted from a misconfiguration by Irregular, an AI security testing and red-teaming firm, which inadvertently granted the model internet access during an evaluation.
Meta told Reuters in a statement that the model “exploited a security vulnerability in a third-party service” in a way similar to previously reported incidents involving other AI companies.
The incident marks the latest example of an advanced AI agent emerging as a cybersecurity risk in its own right. It has also raised questions about where liability should rest—whether with the companies developing the AI agents or those responsible for designing the testing sandboxes intended to contain them.
Anthropic and OpenAI Incidents Highlight Growing AI Safety Risks#
Meta’s AI security incident comes just one week after Anthropic disclosed that its models gained internet access and hacked an external company because of a configuration error in Irregular’s testing environment.
In a blog post published on July 30, Anthropic said it identified three incidents out of 141,006 evaluation runs in which a Claude model accessed the internet during testing and later gained unauthorized access to systems belonging to three separate organizations.
Stay in the loop
Get crypto news before the market moves
Join thousands of investors who read our daily briefing.
No spam. Unsubscribe anytime.
All three incidents occurred within, or during interactions with, Irregular’s evaluation environment and involved a misconfiguration that left the systems accessed by Claude connected to the live internet.
In July, AI agents developed by OpenAI escaped their offline sandbox and hacked Hugging Face while attempting to cheat on a cybersecurity benchmark test.
Charles Guillemet, chief technology officer of Ledger, described the latest incident as “marketing theatre.”
“Having a model ‘go rogue’ has become the latest AI PR stunt,” he said on Wednesday.
“If your model isn’t escaping sandboxes, ‘hacking’ companies, or pulling off some headline-grabbing exploit, apparently you’re falling behind… The industry doesn’t need bigger stunts; it needs greater trust.”



