
The global reaction when OpenAI’s agents escaped a test environment and broke into Hugging Face’s systems in July was that the machines had become frighteningly capable. Dominic White’s reaction was that OpenAI’s security was frighteningly poor.
His critique comes as OpenAI deals with yet another escape. On 25 September, the company disclosed that an agent in training had reached an outside chatbot on 20 September through a gap in its sandbox’s internet restrictions – the first such breakout since it hardened its test environments after the Hugging Face attack. It has again paused training and tool-using work on its most capable models.
White, MD for South Africa at Orange Cyberdefense and a 16-year veteran of SensePost, the group’s ethical hacking arm, told TechCentral the sophistication narrative tells only half the story.
“When I first heard about this, it was portrayed as sophisticated. That’s not anything new in our industry. Any hack has been the work of a ‘sophisticated’ threat actor according to PR teams for as long as it’s been going. Not always been true.”
Parts of the attack chain were genuinely sophisticated, he said. But when OpenAI presented its account at the Black Hat security conference in Las Vegas in August, “what fascinated me is that the initial attack was so mundane”.
The agents sat in a sandbox – an isolated test environment – that could reach an internal server running Artifactory, JFrog’s package management software, which they used to fetch software packages. OpenAI later found previously unknown vulnerabilities in Artifactory, but the initial opening was simple.
‘Really embarrassing’
“The initial vulnerability was that all the agents had the same shared credential and could just write files to the web server with HTTP PUT requests.” In other words, every agent used the same login and could upload files to the server at will. White called that “really embarrassing”.
“If an organisation gave all of their customers the same password, and then customers could just write files to a web server or delete it, we go, that’s negligence.”
The initial vulnerabilities were “quite frighteningly weak”, he said, which made him question whether OpenAI understood its duty of care.
His second objection is that, rather than locking down after the first breach, “they let it happen again”, this time through an unauthenticated WebDAV request – a file-management command the server accepted without any login. OpenAI’s own timeline bears this out. On 8 July it rebuilt Artifactory, revoked the agents’ credentials and restarted its cybersecurity tests; the same day, agents rebuilt their covert message board using unauthenticated WebDAV requests.
His third is that the agents ran for long periods without guardrails. “They knew they were running long-term agents here without cyber guardrails in place. So they kept it alive for a long time without [proper] monitoring. They’re just like, well, we’ve got a sandbox. Except, it was a weak sandbox.”

OpenAI has conceded much of this. In its incident report, published on 26 August, it said the chain-of-thought monitors it uses to watch its models’ reasoning were not running on these evaluations, and that an internal team had seen the agents’ message board and unauthorised internet access as early as late May – warning signs that, in hindsight, should have prompted an earlier response. The company said the incident “should never have occurred”.
White also pointed to independent analysis showing how readily the agents crossed the line. An investigation by AI safety researchers at METR and Redwood Research found that more than 90% of the 533 agents active on the agents’ message board during the period it examined joined the attack on Hugging Face, many of them recognising it was out of scope. “If it was a swarm of humans, there would be a lot more debate about whether this is ethical, whether this is the right thing to do. But that didn’t exist here.”
His verdict: “You’ve got poorly aligned, long-running, offensive agents running in wet paper bag sandboxes. I think the big problem here is the lab might not be doing their job.”
OpenAI has not released the agents’ full chain-of-thought records – their step-by-step reasoning – White noted, limiting outside scrutiny. It has published excerpts, and gave the METR and Redwood researchers access to about 1 300 raw transcripts on its premises.
Adam Ely, cybersecurity firm Check Point’s GM of AI security, who joined the interview for an upcoming episode of the TechCentral Show, was slower to pile on criticism but did not disagree. He called the incident “one of the most fascinating kind of events that I think we’ve seen in cybersecurity for a long time”.
Individually, he said, the vulnerabilities and the agents’ tactics were well understood. What stood out was “the misalignment of the agent to the human intent” – the gap between what the agents did and what their operators wanted – and the scale at which they organised themselves.
Ely also noted that one of the presenters at OpenAI’s Black Hat talk is “probably one of the smartest people that I’ve known throughout my entire career”.
Every dependency is a door
“When you look at the security teams, these are knowledgeable people. This wasn’t a fly-by-night, just kind of thrown together set of organisations.”
Ely’s own researchers had seen similar behaviour before, in environments where the open question was whether it could happen to a serious operation. “Would a well-resourced team, a responsible company run into the same problems? Or were we only seeing this in teams that were just starting out and they hadn’t got going enough yet with their protections?” In his view, the Hugging Face incident answered that question.
Test environments still need to pull packages and reach internal services, and every dependency is a door. “Anytime that you have to rely on an external service, maybe it’s a DNS service, maybe it’s a package manager, a vulnerability, a misconfiguration, that’s the gateway out or the gateway in.”
OpenAI’s 20 September escape fits that description. The agent reached the outside chatbot through inadequate filtering of DNS, the system that translates web addresses into network locations.
What was new, Ely said, was seeing long-understood weaknesses exploited at massive scale, without humans. No amount of resourcing removes the risk: “No matter what you resource and how you design, the fact is there’s some level of probability that agents are going to go rogue.”

The lesson Ely draws is not that OpenAI was uniquely careless, but that ordinary discipline now carries more weight: “The basics really matter in the infrastructure and the design and the incident response and the detection. Those basics now are even more important because we have to move faster, we have to see things faster.”
He echoed White on consequences. “If these were humans doing these things, we would instantly say these [actions] are unacceptable. There would be lawsuits. But for some reason AI is getting a little bit of a [pass] right now.”
The full podcast interview with White and Ely will be published on TechCentral this week. — © 2026 NewsCentral Media





