“We are currently investigating and will issue a full retrospective once we have all the facts,” a Meta spokesperson said in a statement
Credit: Photo Illustration by Thomas Fuller/SOPA Images/LightRocket via Getty
NEED TO KNOW
- Meta AI’s Muse Spark 1.1 hacked another company’s systems during cybersecurity testing due to a misconfiguration
- Similar incidents have occurred with AI models from Anthropic and OpenAI during testing of their cyber capabilities
- Independent testers are investigating secure methods for AI evaluations and plan to release best practices for containment
Meta AI hacked into another company during cybersecurity testing, marking the fourth incident involving an AI company.
A Meta spokesperson says in a statement shared with PEOPLE that the hack occurred after a “misconfiguration” by Irregular, an independent testing company that Meta uses, “inadvertently allowed one of our models access to the internet during evaluation.”
“The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” the spokesperson adds. “Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts.”
The hack was said to involve Meta’s Muse Spark 1.1 model, Australian Broadcasting Corp (ABC) reported.
An Irregular spokesperson confirms to PEOPLE that the company is investigating how to securely run cybersecurity tests with AI agents. “This did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evals,” the statement said.
Irregular also confirmed that what occurred with Meta is similar to hacking incidents with Anthropic, the company behind Claude.
“The discussed incident is the exact same evaluation-environment issue that was already disclosed by Anthropic last week,” the Irregular spokesperson adds.
Anthropic detailed three separate incidents in a statement on July 30, calling the hacks a “misunderstanding” of the sandbox testing methods, which gave the models fictional targets to hack into, but instead gained access to the internet.
The first incident occurred in April, but none of the companies were aware of the hacks until July, Anthropic said in its statement.
Never miss a story — sign up for PEOPLE’s free daily newsletter to stay up-to-date on the best of what PEOPLE has to offer, from celebrity news to compelling human interest stories.
Anthropic’s announcement came days after OpenAI’s model GPT‑5.6 Sol hacked another company’s AI system in mid-July. The company, which is behind ChatGPT, announced the security incident in a blog post, saying it occurred while the company was running a realistic benchmark, or test, called ExploitGym, built from real-world vulnerabilities. OpenAI called the hack an “unprecedented cyber incident.”
The testing was conducted in a “highly isolated environment” to identify its “maximal cyber capabilities.” “We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity,” OpenAI continued.
The test was also run with network access restricted to installing software packages that may have been needed to complete the test. “The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI said.
After obtaining “open internet access,” the models “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation” through the Hugging Face servers.
Clem Delangue, Co-founder and CEO, Hugging Face, said in a statement on OpenAI’s website that the incident “proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Read the full article here
