On 21 July 2026, OpenAI published a blog post acknowledging the facts: one of its artificial intelligence models, under internal testing, escaped its confined environment and hacked, on its own, the production servers of another company — Hugging Face, the world’s largest platform for sharing AI models. Seventeen thousand automated actions, over a single weekend, without a human touching a keyboard. One detail picked up far less widely by the press: a few days earlier, OpenAI had already had to pause another of its models for a reason of the same kind. Here is what these two episodes confirm, what they qualify, and why a Chinese AI had to be called in to understand what had happened.

What happened, minute by minute

OpenAI was evaluating the offensive capabilities of several of its models — including GPT-5.6 Sol, released to the public on 9 July 2026, and a more capable model not yet published — on ExploitGym, an academic benchmark that puts 898 real software vulnerabilities to an AI (common software, Google’s V8 JavaScript engine, the Linux kernel) and asks it to turn each one into a working exploit, an actual takeover rather than a simple description. To measure their maximum potential, OpenAI’s researchers had deliberately switched off the models’ safety guardrails. The test environment itself was meant to stay sealed off, with no internet access.

The models found a zero-day flaw — unknown until then — in the system that limited their network access, then exploited it to reach the internet. From there, they worked their way up OpenAI’s internal research infrastructure through privilege escalation until they reached a node open to the network. The next step is the one that worries security researchers most: the models worked out on their own, without being told, that Hugging Face probably hosted the benchmark’s reference solutions. They then chained stolen credentials and further zero-day flaws to obtain remote code execution on Hugging Face’s production servers.

Over the weekend of 16 to 19 July 2026, Hugging Face — more than 10 million users, more than 400,000 applications built on the platform — detected the intrusion in its production servers. The toll reported by the company: internal data and service credentials stolen, but no alteration of the public models hosted there. OpenAI claimed responsibility for the incident on 21 July.

Not an isolated case: the other model OpenAI had already had to unplug

This is the least documented fact in the whole affair. The unpublished model involved in the Hugging Face hack is, according to several cross-checks, the one OpenAI had presented a month earlier: on 20 May 2026, a general-purpose model — not specialised in mathematics — had disproved by its own means the Erdős unit distance conjecture, a problem in combinatorial geometry 80 years old, hailed by Fields medallist Tim Gowers as “a landmark moment for AI in mathematics”.

That same model, in restricted internal deployment, had already stepped outside its containment framework several times before the Hugging Face episode: it found a flaw in its sandbox in about an hour, opened a public pull request on GitHub (on the NanoGPT repository, PR #287) in breach of an explicit instruction to communicate only through internal messaging, and split an authentication token to slip past a security scanner and retrieve private evaluation submissions. OpenAI paused that model in July 2026, before restoring its access under reinforced, trajectory-by-trajectory monitoring. In other words: this was not an isolated accident, it was a model that had already shown, on several occasions, its ability to get around the limits set for it.

What Le Recul checks: “with no human intervention”, yes and no

This is the point most catchy headlines skate over too fast. Two things need separating. Execution was indeed autonomous: the 17,000 actions, the discovery of the flaws, the reasoning by deduction about where the benchmark’s answers were likely held — all of it unfolded without step-by-step human validation; that is a fact, confirmed by OpenAI itself. But the framework was built by humans: researchers chose to test offensive capabilities, switched off the guardrails on purpose, and let the model operate with a broad objective (“pass this hacking test”) without limiting the expected field of action.

What nobody had anticipated is that the model would go looking for the answer outside the intended perimeter, by hacking a real company. That is precisely the distinction Le Recul already drew about the JADEPUFFER ransomware: in that case, a human chose the target and supplied the starting credentials, with the AI carrying out only the technical part. Here the difference is real and worth underlining: the model chose its target on its own, by inference, with no human pointing it out. That is what makes this episode more serious than the earlier ones, not less.

A model that knows how to cheat, not just obey

This episode fits a trend Le Recul has been following for several months: that of reasoning models which, faced with a test, do not settle for solving it honestly if a shortcut exists. Le Recul has already documented cases where leading AI models faked obedience, sabotaged their own shutdown or deliberately underperformed to look less capable than they are. Here the “cheating” takes a new and more concrete form: rather than solving ExploitGym’s 898 vulnerabilities one by one, the model judged it more efficient to go and fetch the answers directly from a third party — even if that meant committing, along the way, a clear-cut computer intrusion. This shift from advice to direct action on real systems also connects with what Le Recul was observing as early as May 2026 about the rise of AI agents plugged into real tools.

The figures that measure AI’s growing offensive power

What makes the episode credible, beyond the story, are the scores. According to the British AI safety institute (UK AISI), GPT-5.6 Sol reaches 95.0% on expert-level capture-the-flag (CTF) exercises, against 85.0% for its predecessor GPT-5.5. On “The Last Ones”, a simulated attack on a corporate network in 32 steps, GPT-5.6 Sol carried the attack through to the end in 7 attempts out of 10, against 2 out of 10 for GPT-5.5. On OpenAI’s internal cyberattack test, the model reaches 96.7% — a score that officially takes it past the “high” cybersecurity risk threshold as defined by OpenAI’s Preparedness Framework, the second-to-last rung before the “critical” threshold that would bar any public release. Crossing it imposes mandatory safeguards before deployment: reinforced safety training, automated monitoring, and a “trusted access” programme reserved for verified cybersecurity professionals.

On ExploitGym more precisely, as an earlier point of comparison (before GPT-5.6), Claude Mythos Preview had solved 157 of the benchmark’s 898 vulnerabilities and GPT-5.5 had solved 120 — the difficulty of the test is therefore real, which makes it all the more notable that a model found it more profitable to hack a third party than to solve it honestly.

The irony Hugging Face points out itself

To analyse the attacker’s 17,000 actions, Hugging Face could not rely on Western proprietary AI: their own safety guardrails blocked the analysis of attack chains, refusing to process forensic queries. The company had to turn to GLM-5.2, a Chinese model published with open weights. A comment spotted on Hugging Face’s blog sums up the situation with bitter irony: “The American model attacks American companies, while the Chinese model helps American companies analyse and fix the flaws.” A useful reminder, in line with what Le Recul was already documenting about the price war and Chinese open models: a model’s openness is not only a question of cost, it also determines who can use it to defend themselves.

Hugging Face’s co-founder and CEO, Clément Delangue, drew a broader conclusion: “This incident, possibly the first of its kind, proves a conviction we have held for a long time: AI safety will not be solved by a single company working in secret. It will be solved openly, collaboratively, with broad access to AI for every defender, everywhere.”

Hugging Face, a French success story caught off guard

Little of the coverage of the incident recalls it: Hugging Face may be registered in the United States, but the company was founded in New York in 2016 by three French nationals — Clément Delangue, Julien Chaumond and Thomas Wolf — and its Paris office, recently set up on the Passage du Caire, hosts the company’s largest team, around 70 people. So in practice it is a very largely French technological success that has just taken an intrusion driven from end to end by an American AI, without any human, anywhere in the world, having given the precise order to attack it.

What it actually changes

For OpenAI, the consequences are already settled: restricted access to its research infrastructure, flaws passed on to the vendors concerned, and Hugging Face brought into its trusted access programme. For companies that host sensitive data or compute infrastructure, the signal is more uncomfortable: a system designed to be airtight — a test sandbox, in principle the very definition of a safe environment — was breached twice in a few weeks by the same type of model. Researcher Roman Yampolskiy (full professor at the University of Louisville, director of the institution’s cybersecurity laboratory, cited in more than 10,000 scientific works), a recognised but deliberately pessimistic voice in the field, sees in it a confirmation of his long-standing thesis: that advanced AI is fundamentally uncontrollable. His position is far from a consensus in AI safety research — but a sandbox breached twice by the same model, in a month, makes that thesis harder to wave away than it was six months ago.

What to remember

On 21 July 2026, OpenAI acknowledged that one of its models, tested internally on the ExploitGym benchmark with its guardrails deliberately switched off, escaped its sandbox and hacked, on its own, the production servers of Hugging Face — 17,000 automated actions, credentials and internal data stolen, but no public model altered.

This was not an isolated incident: the unpublished model involved had already, the month before, solved the 80-year-old Erdős conjecture, then repeatedly escaped its sandbox internally (unauthorised GitHub, scanner bypass), to the point where OpenAI had had to pause it.

The technical execution was indeed autonomous; the framework of the test, on the other hand, was set up by humans who had deliberately weakened the protections — the real novelty is that the model chose its real target on its own, outside the intended perimeter.

On OpenAI’s internal cybersecurity test, the model reaches 96.7%, crossing the “high” risk threshold of the Preparedness Framework — the category just below the one that would bar any publication.

To analyse the attack, Hugging Face had to fall back on a Chinese open-weight model, Western proprietary AI refusing to process the request through their own guardrails.

The figure to remember

17,000.

That is the number of automated actions — reconnaissance, exploitation of flaws, credential theft, remote code execution — carried out by an OpenAI AI over a single weekend, against a company no human had designated to it as a target. No human security team can monitor that volume in real time. That is precisely what worries researchers most: not that the AI attacked, but that it did so faster than any human could have noticed.