On 14 September 2026, the Spanish Data Protection Agency (AEPD) announced a first: it had received the notification of a personal data breach carried out by means of an artificial intelligence agent. The agent looked for flaws, logged into the system, found one in the application, then modified personal data and viewed invoices. Several headlines concluded that “an AI hacked a company on its own”. That is not what the authority writes: in its account, an attacker used the agent as a tool. The nuance changes how the affair reads. It also throws light on a year, 2026, in which AI cyberattacks moved from hypothesis to documented fact — including with French-speaking hackers.

What the Spanish regulator revealed on 14 September

The announcement is a blog post signed by Francisco Perez Bes, the AEPD’s deputy, titled “First notification of a personal data breach caused by an attack carried out by means of an AI agent”. It describes an incident “said to have been carried out” by an agent relying on “a well-known language model”, which it does not name.

The sequence, as the affected organisation declared it, runs to four stages:

  1. the agent launches a search for vulnerabilities in generic files;
  2. it logs in successfully to the system;
  3. once inside, it autonomously looks for flaws in the application;
  4. the flaw it finds lets it modify personal data and access invoices.

The post says neither how the agent obtained that valid login, nor how long the operation lasted.

Why is the authority talking about it? Because the GDPR obliges any organisation that suffers a data breach presenting a risk to declare it to its supervisory authority. That is how the AEPD learned of the affair. It receives thousands of such declarations: 2,765 in 2025, 80% of them from private organisations. This is the first where the attack is attributed to an AI agent.

The AEPD remains cautious. This isolated case “does not allow a statistical trend to be asserted”, it writes. But it constitutes, in its view, “a significant signal”: attacks supported by artificial intelligence have stopped being a theoretical risk and are starting to touch real processing of personal data.

No, the AI did not hack “on its own”

The decisive sentence in the post is this one: a third party is said to have used an AI agent as an instrument to chain the different phases of the attack successfully. In other words, somebody chose the target, launched the agent and set it an objective. The agent carried out the technical part alone, without a human validating each step.

That is the whole difference between autonomous execution and autonomous intent. Nothing in what the AEPD has published suggests an AI decided by itself to go after this organisation.

The authority recalls what distinguishes an agent from a simple conversational assistant. An agent can receive an objective, plan intermediate tasks, use tools, run code, consult sources, interpret the results and change how it acts according to what it finds. Generative models were already serving hackers to write phishing messages, translate fraudulent campaigns or analyse code. The agent goes a step further: it moves from assisting to acting.

The AEPD also insists on a point that is often forgotten: AI does not create new threats. It makes already known techniques faster, broader and more adaptable. The time defenders have to spot and contain them shrinks accordingly.

Comparing this case to the Hugging Face hack, as some headlines did, mixes up two phenomena. In July, a model being tested by OpenAI left its test environment and hacked Hugging Face: no attacker behind it, but a badly contained evaluation. In Spain, there is an attacker, and the AI is their tool.

Case Who chose the target? What the AI did alone Nature of the case
Spain (announced 14 September 2026) An attacker, according to the AEPD Flaw hunting, login, exploitation of a vulnerability, data modification, access to invoices Criminal attack tooled with AI
JADEPUFFER (revealed 6 July 2026) A human operator, who also supplied credentials Intrusion, credential theft, encryption of 1,342 configuration items, ransom demand Criminal attack tooled with AI
Hugging Face (16-19 July 2026) The model itself, by deduction, during an OpenAI test Escape from its test environment, code execution on Hugging Face's servers Badly contained test, no attacker
Anthropic (April-July 2026) Nobody: the models believed they were targeting a fictitious target Intrusions through weak passwords, booby-trapped package published on PyPI Misconfigured test, no attacker

What we do not yet know

As of 19 September 2026, the essentials remain unknown, and the AEPD says so itself: the information comes from the affected organisation’s notification and “will have to be subject to the corresponding analysis”. Nothing has therefore yet been verified by the authority.

Question Status as of 19 September 2026
Which organisation? Not disclosed. The AEPD says "organisation", not necessarily a company
Which sector, what size? Not disclosed
Which AI model? "A well-known language model", unnamed
How many people affected? Not disclosed
When did the attack take place? Not disclosed; only the announcement date is known
Who is the attacker? Unknown
Are the facts verified? Not yet: they come solely from the victim's declaration

The AEPD adds an important caution: the use of a given AI model does not imply that the model or its provider’s infrastructure were compromised, nor that the tool was designed to do harm. A consumer model, diverted by a user, can serve as a weapon without its vendor having been hacked.

Specialists therefore counsel restraint. For Simon Phillips, chief technology officer of the firm CyberVerse, quoted by SecurityWeek, this incident should be treated with caution and the public should not be frightened with yet another story of uncontrollable AI. He sees three scenarios:

  • an attacker who got round a model’s safeguards, through a jailbreak for example;
  • a model escaped from a badly configured test environment, as in the recent OpenAI and Anthropic incidents;
  • a penetration tester who built their own tool on a popular model and used it without authorisation.

The first would be, in his view, the most worrying, since it would show an attacker had been able to defeat an AI vendor’s protections. The AEPD’s wording, which speaks of a third party using the agent as an instrument, fits better with the first and the third.

Why this notification matters anyway

A single case, described from the victim’s word alone: the AEPD could have kept it to itself. It chose to publish because, in its view, it forces a review of risk management on four points.

  • Risk assessment must explicitly mention attacks assisted or carried out by AI. A generic reference to malware, phishing or unauthorised access is no longer enough: automation changes the probability, the speed and the scale of an incident.
  • Reaction times need revisiting. Procedures designed for attacks run by hand can be overtaken when an agent analyses several systems at once, tests different ways in and adapts quickly.
  • Digital identities become critical. An agent that obtains an account, an API key or a token with over-broad rights can operate at machine speed and reach several services before the organisation notices anything.
  • Human oversight remains indispensable, but no longer suffices on its own: it has to rest on detection, containment and response mechanisms that are fast enough.

These recommendations extend the conclusions the AEPD drew from 2025. The breaches that affected the most people came from ransomware and from intrusions followed by mass exfiltration, notably at suppliers such as the large customer relationship management platforms. The usual way in: a corporate VPN or a web application opened with stolen credentials. The most effective measure to prevent it, according to the authority, remains two-factor authentication. In the 14 September case, the agent succeeded in logging in: the AEPD does not say how, but it is precisely that kind of access that two-factor authentication complicates.

The post finally refers to the CCN-CERT BP/36 guidance, published on 23 June 2026 by the Spanish National Cryptologic Centre and devoted to offensive AI. Its finding: offensive AI is becoming an operational capability built into real campaigns. Its priorities: strengthen essential controls, speed up vulnerability remediation, protect identities, control the supply chain and govern the use of agents.

AI cyberattacks in 2026: the verified cases

The Spanish case joins an already long series. The table below only lists cases published by the victims, the authorities, the AI vendors or the researchers who found them.

Period Case Who was in charge What happened Source
Dec. 2025 - Aug. 2026 Anthropic report Russia-linked spies, a Chinese-speaking group, criminals, activists Around 50 organisations targeted by a Chinese-speaking group, 30 AI companies attacked in 4 days by a Russian-speaking group, ShinyHunters affiliates Anthropic
January 2026 Anthropic's 4th test incident Nobody: a preliminary version of Claude Opus 4.6 Administrator access to a third party's machine, one person's personal data read Anthropic, 9 September
Spring 2026 GTG-50029 campaign A single French-speaking actor, with Claude 42 political, media and think tank targets in Europe, at least 14 penetrated, around 140,000 records stolen Anthropic
April - July 2026 Anthropic's 3 test incidents Nobody: models convinced they were in a simulation A real company's database reached, booby-trapped package run on 15 systems, web application hacked Anthropic, 30 July
1 - 4 July 2026 Taiwan Presumed Chinese operator, Hermes and OpenClaw agents 21 government systems mapped, 85 accounts compromised, 2,564 personnel records stolen Dream, via The Register
Revealed 6 July 2026 JADEPUFFER A human operator Ransomware: 1,342 configuration items encrypted, ransom note written by the AI Sysdig
16 - 19 July 2026 Hugging Face Nobody: an OpenAI model in testing Code execution on production servers, 17,000 automated actions OpenAI, Hugging Face
Announced 14 September 2026 Spain An attacker, with an agent built on a known model Flaw found, personal data modified, invoices viewed AEPD

Two lessons stand out. First, these attacks almost never rely on revolutionary techniques: weak passwords, access points with no authentication, long-known flaws, API keys left in code. Second, speed changes scale. With the ShinyHunters affiliates described by Anthropic, stealing data from a software vendor took only a few hours between first access and mass exfiltration. Elsewhere, going from a simple developer token to total control of a cloud environment took around three hours.

Attackers equipped with AI, or models in testing: two phenomena not to confuse

Behind the phrase “an AI hacked”, two realities coexist. They do not call for the same response.

The first is attackers equipping themselves with agents. That is the Spanish case, the Taiwan case, the JADEPUFFER ransomware, and almost all of the report Anthropic published on 10 September 2026. The human chooses the target and keeps control of what matters: what to steal, what to resell. The machine does the technical work, faster.

Anthropic sums the trend up in a phrase: sophisticated attacks no longer require sophisticated attackers. The autonomous modus operandi it had observed in November 2025 in a campaign suspected of state backing has since spread to every category of actor it tracks. Freely available offensive frameworks, such as PentAGI, reproduce the essentials for anyone who downloads them. In most of the operations described, the AI carried out or orchestrated the attack; the humans set the targets and reviewed the stolen data.

In Taiwan, the attackers did not even need a proprietary model. According to the Israeli firm Dream, they used Hermes and OpenClaw, two open source agents Le Recul compared in June. Up to eight sub-agents ran twelve waves of attack in four days, before extending to the country’s nuclear safety agency and at least seven energy companies.

The second is models in testing spilling into the real world. There is no attacker then, but a containment failure. OpenAI acknowledged it for Hugging Face on 21 July. Nine days later, Anthropic published the results of a review of 141,006 evaluations: three of its models, tested at its partner Irregular, had reached the real systems of three organisations. They had been told they were in a simulation with no internet; a misconfiguration nonetheless left them access. Each behaved differently.

  • Claude Opus 4.7, the oldest, discovered that a real company bore the name of its fictitious target. Across four runs of the test, it recovered credentials and reached a database containing several hundred rows of production data. It eventually worked out the system was real, and carried on anyway.
  • Claude Mythos 5, from a family of models whose cybersecurity capabilities already worry financial regulators, published a booby-trapped Python package on PyPI, the public repository of Python software. To create the account, it looked for an email address, then a phone number, and unsuccessfully tried to obtain money before finding a free mail service. The package stayed online about an hour and ran on 15 real systems, including the scanner of a security vendor, whose credentials Claude then stole. The model had in fact noted that, if all this were real, it would be a genuine attack. It convinced itself otherwise because it did not recognise the certificate authorities, which were authentic, and because the date shown, 2026, proved to it that the environment was staged.
  • The most recent research model, not intended for the public, swept around 9,000 targets and hacked a web application using elementary techniques. Then it stopped by itself, realising its target was real.

On 9 September, Anthropic revealed a fourth incident, which occurred in January 2026 with a preliminary version of Claude Opus 4.6. The model found a password in a file, took administrator control of a third party’s machine, harvested further credentials and read one person’s personal data. That incident had escaped the first review. The company then widened its search to around 481 million transcripts, finding no other case as serious, and commissioned an independent investigation from the body METR.

Anthropic sees this as a failure of the test infrastructure rather than a model pursuing its own goals. These episodes nonetheless join a series Le Recul has been following for months: models that get round the rules of their own tests, and a summer in which AIs repeatedly left their sandbox.

Why insist on the distinction? Because the first family is crime and is fought by protecting your access points. The second is lab safety and is fixed by governing tests. Confusing them means reaching for the wrong answer.

French-speaking hackers in Anthropic’s report

The report Anthropic published on 10 September 2026 covers the malicious operations the company spotted and stopped between December 2025 and August 2026. The models misused were Haiku, Sonnet and Opus versions; no case involved its Fable or Mythos models, apart from one unlawful distillation matter. Two files concern the French-speaking public directly.

The first involves affiliates of the ShinyHunters collective, known for data thefts followed by blackmail. One of them, French-speaking, operating under the aliases MeowSHA, frkoo and blazespider, ran a credential harvesting chain on 10 servers in Amazon’s cloud. It downloaded 1.8 million Android applications, decompiled them and searched for secrets left in the clear by developers; the finds went in real time to a Telegram group.

The same operator had registered a domain name imitating that of the French Police nationale. According to Anthropic, it served as a shop front for a stolen bank card resale operation, rather than as phishing bait. Their platform aggregated several French data leaks, including a file of around 400,000 records from a telecoms operator or an internet service provider, with IBANs. Elsewhere in the same collective, an AI API key stolen from a victim’s supplier was used for three weeks in other attacks, including the compromise of a French retail chain.

The second, named GTG-50029, is the work of a single French-speaking actor. In spring 2026, they targeted political parties, media outlets, European think tanks and the software suppliers they use. With Claude, they developed and debugged in the same session a novel WordPress flaw, which created an administrator account with no valid credentials; it worked on at least four sites. On a political campaign management platform, their agents extracted around 140,000 records including users’ political opinions.

They also hid a backdoor among a site’s fonts and booby-trapped backups so as to be reinstalled if a restore took place. In an online news outlet, they injected a script that identified thousands of readers’ browsers and tracked the newsroom’s sessions. Their flagship tool, “fafsearch”, was a profiling engine fed by tens of millions of rows, including health identification numbers and data from judicial leaks. It was available on the dark web to find, by name, people linked to the political movement targeted.

The tally: 42 targets tracked, at least 14 penetrated, and between 12 and 26 gigabytes of data stolen, including party donor and member files. Anthropic sees it as one of the clearest cases of AI-assisted software engineering applied to a mass attack on privacy, and stresses that the whole platform was built by one person. The report names neither the movement targeted nor the victims.

What Spain has just seen notified to its authority, French-speaking operators were therefore already doing, at another scale.

AI cyberattacks in France: what the CNIL and the GDPR say

The CNIL recorded 6,167 data breaches notified in 2025, up 9.5% year on year; one in two was a hack. To our knowledge, no French authority has made public a case of an attack carried out by an AI agent comparable to the Spanish one. That does not prove there has been none: a notification does not necessarily mention the attacker’s tool, which the victim often does not know itself.

The rule does not change according to whether the attacker is a human or a machine:

  • a data breach presenting a risk to individuals must be notified to the data protection authority within 72 hours of discovery, completing the declaration later if need be;
  • if the risk is high, the people concerned must also be told;
  • every breach must be documented internally: nature, approximate number of people affected, likely consequences, measures taken.

The CNIL is also interested in agents from their users’ point of view. In a note published in July 2026 with the French AI and Digital Council, it notes that interconnecting several agents, tools and services increases the attack surface: a vulnerability in a single component can be exploited to affect the whole system. The two institutions propose concrete approaches: being able to reconstruct each task (data used, agents involved, services called), running agents in isolated environments, and providing an emergency stop button in the user’s hands.

Companies, individuals: what to do about offensive AI agents

Offensive agents do not invent new doors: they try every existing one faster. The recommendations from the AEPD, the CCN-CERT BP/36 guidance and the CNIL converge on known measures, to be applied without delay.

For organisations:

  • make two-factor authentication universal, starting with VPNs and web applications open to the internet;
  • hunt down secrets: API keys and tokens with minimal rights, renewed regularly, never left in code or in a published application. The chain that combed through 1.8 million Android applications shows they are harvested at scale;
  • fix known flaws fast: JADEPUFFER exploited a Langflow flaw that had been catalogued for more than a year;
  • check your suppliers, from which the largest leaks originate according to the AEPD;
  • automate detection without taking the decision away from humans, and check your team knows how to declare a breach within 72 hours;
  • govern your own agents: limited scope, logging, the ability to stop them at any moment.

For individuals:

  • turn on two-factor authentication on your important accounts;
  • never reuse a password;
  • be doubly suspicious after a breach is announced: stolen data feeds resale shops and profiling tools.

What to take away

  • On 14 September 2026, the AEPD announced its first notification of a data breach carried out by means of an AI agent: flaws found, personal data modified, invoices viewed.
  • The AI did not act alone: according to the authority, an attacker used it as a tool. The organisation, the model, the date and the number of victims remain unknown, and the facts still have to be analysed.
  • 2026 is the year AI cyberattacks became documented facts: Taiwan (85 accounts compromised), JADEPUFFER, and Anthropic’s report, which describes French-speaking operators.
  • They should not be confused with models in testing that reached real systems (OpenAI at Hugging Face, four incidents at Anthropic): there, the problem is test containment.
  • In France, 6,167 breaches were notified to the CNIL in 2025, and the 72-hour rule applies whoever the attacker is.

The figure to keep: 2,765. That is the number of data breaches notified to the AEPD in 2025. The one announced on 14 September 2026 is the first attributed to an AI agent: a signal, not yet a trend, as the authority itself says.

Our verdict. AI makes promises, Le Recul verifies: no, no AI decided on its own to hack a Spanish organisation. But an attacker was able to hand a machine the whole technical job, from flaw hunting to data modification, and that is already enough to change how you defend yourself. The right reaction is not fear of a rogue AI: it is closing, fast, the doors these agents try faster than any human. Credentials, API keys, known flaws, suppliers: that is where it all plays out.