On Thursday 20 August 2026, a new line appeared in the catalogue of OpenRouter, the platform that brings hundreds of artificial intelligence models from around the world together behind a single access point. It was called stealth/ox-alpha.
No vendor name. No price. No press release. Just a spec sheet, an API address, and a two-line description: “reasoning model designed for code, long-horizon agentic work and production workloads”.
It accepted 1,048,576 tokens of context — a million, enough to swallow an entire code base. It read text, images and video. It called tools. And it was free.
Within hours, the screenshots started circulating. Patrick Collison, Stripe’s chief executive, publicly called it “very impressive”. The anonymous provider announced a capacity of a hundred trillion tokens a day. On French-speaking networks, a video summed the affair up in nineteen seconds: a terrific AI, free, whose author nobody knows, and — this is the sentence that interests us — “it says zero data retention”.
It did not say that. It said the opposite.
Six days, a million tokens, and no name
Here is the sequence, as it unfolded.
| Date | What happened |
|---|---|
| 14 August 2026 | Z.ai launches GLM-5.3, its flagship model, but announces it will hold the weights back for about two weeks, pending a safety evaluation. |
| 20 August 2026 | Ox Alpha appears on OpenRouter. Free, a million tokens of context, anonymous provider. |
| 21-25 August | The model goes viral. Guesses split between Z.ai and Microsoft's MAI family; tokeniser analysis leans towards Microsoft. |
| 26 August | Z.ai lifts the mask: Ox Alpha was GLM-5.3-Flash, 320 billion parameters, MIT licence. |
| 27 August | Zhipu AI stock closes up about 12% in Hong Kong, at 1,160 Hong Kong dollars. |
| 28 August | Z.ai finally publishes the GLM-5.3 weights — under an in-house licence, not under MIT. |
Six days of complete fog, a reveal, and a model that goes from anonymous curiosity to financial event. In between, thousands of developers plugged that model into their real work, because it was free and it was good.
It is that six-day window that deserves a close look, far more than the spec sheet.
The sentence the video misread
The model’s page on OpenRouter carried a data policy note. Here it is in full:
“Prompts and completions for this model were retained by the provider and are not used for training; any other use is governed by the terms applicable to stealth models.”
Read it twice. It says two distinct things, and the viral video kept only one.
It says first that prompts and completions are retained. That is the exact opposite of “zero data retention”. The confusion comes from a widespread shortcut: “not used for training” was read as “not kept”. These are two completely different guarantees. The first is about a use. The second is about possession.
The distinction is not a legal detail. Data that trains no model but sits on a server is still data somebody holds, that a third party can demand, that a breach can expose, and that a court order can go and fetch. The relevant question was never “does it train the AI?” but “who has my data, where, for how long, and under which law?”.
For six days, the answer to those four questions was: nobody knows.
The second document, the one nobody opens
There is something more awkward, and it went almost unnoticed.
The note quoted above refers explicitly to another text: the terms of OpenRouter’s Stealth programme. And that framework agreement, in its version updated on 6 July 2026, provides that user content may be collected, shared and licensed for model training and improvement purposes.
In other words, the model card and the framework agreement it points to say opposite things on the most sensitive point. Did Ox Alpha benefit from an exception negotiated with its provider? Which of the two texts prevails in the event of conflict? OpenRouter has not publicly settled it.
This is not an accusation. Cascading contracts regularly produce that kind of inconsistency, and nothing indicates abusive use took place. But you have to measure the position it put the user in: they were sending code to an entity they could not name, under a contractual regime the public documents did not allow them to determine.
You cannot audit a company you cannot name. No jurisdiction. No published retention period. No signable processing contract. No verifiable security posture. Nobody to contact if something goes wrong.
Why this is a specifically European problem
An independent developer testing a model on a throwaway project runs no serious risk. A European company plugging the same model into its product does.
The GDPR requires a controller to bind its processor by a written contract identifying the parties, the nature of the processing, the retention period and the security measures. You do not sign that contract with an anonymous party. Nor do you document a transfer outside the European Union when you do not know where the servers are — and nothing on the card indicated a location.
Add the AI Act, with one point that matters. Contrary to what is often written, 2 August 2026 did not trigger the bulk of the regulation: the obligations on high-risk AI were pushed back to 2 December 2027 by the Digital Omnibus, as we set out in our analysis of that deadline. What does apply, on the other hand, are the obligations on general-purpose models — transparency, a summary of training data, respect for copyright — in force since August 2025, with fines reaching 15 million euros or 3% of worldwide revenue. A provider that refuses to name itself cannot, by construction, satisfy any of those transparency obligations, and the company integrating it inherits the hole in its own documentation.
The practical rule fits on one line, and it applies to every stealth preview to come: public code, throwaway projects, synthetic data — never proprietary code, customer data, credentials or content covered by a confidentiality agreement.
Who was behind Ox Alpha: the answer came on 26 August
Through the week, two hypotheses circulated. The first pointed to Z.ai, the Chinese lab formerly known as Zhipu AI, which had already practised anonymous launches. The second pointed to Microsoft’s MAI family, and analysis of the model’s tokeniser seemed to support it. Analyst Andrew Curran summed up the state of the discussion by noting that the more time passed, the less sure observers were of anything.
On 26 August, Z.ai settled it itself: Ox Alpha was GLM-5.3-Flash, “previously previewed as Ox Alpha”, in the words of its announcement.
It is the same lab we were writing about a few days earlier on another subject: Mistral now hosts a Z.ai model on its own servers, GLM-5.2 in that case. In two weeks, the same Chinese company therefore turns up both inside Europe’s sovereign offering and behind the most discussed anonymous model of the summer.
What GLM-5.3-Flash is really worth
The model is not a curiosity. It is a mixture-of-experts architecture of 320 billion parameters in total, of which only 18 billion are activated per token. Forty-five layers, routing over 8 experts out of 288, hybrid attention, natively FP8 weights, a multimodal training corpus of 30,000 billion tokens.
What that produces, as measured by Artificial Analysis, an independent evaluator our own ranking uses as a source:
| Test | GLM-5.3-Flash | Comparison |
|---|---|---|
| Intelligence index (AA v4.1.1) | 57 | Claude Opus 4.8: 57 — identical score |
| Cost per index task | $0.09 | Claude Opus 4.8: $2.03 — 22.6 times more |
| Terminal-Bench 2.1 | 84.3 | Opus 4.8: 85.0 — GPT-5.6 Terra: 87.4 |
| OfficeQA Pro | 62.4 | ahead of Opus 4.8 |
| DeepSWE v1.1 | 63.4 | GLM-5.2: 46.2 — a generation's progress |
| Rank among open-weight models | 4th of 111 | 51st on speed, 16th on cost |
A composite index remains an average, and it should not be turned into a verdict: depending on the test, the gap with the American models widens one way or the other. But the top line is hard to play down. The same composite score. Twenty-two times cheaper.
The price, and what it does to the market
The official price of GLM-5.3-Flash is $0.15 per million input tokens and $0.50 output, with $0.03 for cached input. A 50% discount applies until 9 September 2026. In yuan, that gives 0.8 and 2.8 yuan per million tokens. The flagship GLM-5.3 is billed at $1.40 input and $4.40 output: the Flash costs nearly ten times less than the same vendor’s top of the range.
That positioning is not new: it applies a strategy we documented in detail in why China gives its AI away free. What changes here is the level reached: until now, the trade-off was explicit — cheaper, slightly less good. On this particular index, the trade-off has gone.
The market understood immediately. Zhipu AI is listed in Hong Kong; the stock closed up about 12% at 1,160 Hong Kong dollars the day after the reveal. A week of free evaluation by thousands of enthusiastic developers, followed by a spectacular reveal, translated directly into market capitalisation.
Can you run it at home?
That is the question everybody asks, and the answer deserves precision, because optimistic claims circulate freely.
No, not on a desktop machine. The weights come to around 306 gigabytes in FP8, and the self-hosting documentation published with the model asks for Nvidia Hopper GPUs or newer, with a node of eight GPUs minimum. The supported engines are SGLang, vLLM, TokenSpeed and KTransformers.
The fact that it activates only 18 billion parameters per token makes it fast and cheap to serve — that is the whole point of the architecture. But the full 320 billion have to be loaded into memory. A mixture of experts saves compute, not memory. The nuance is systematically lost in enthusiastic summaries, and it is what separates “a light model” from “a model that is cheap to run at scale”. GLM-5.3-Flash belongs to the second category, not the first.
If your need is really to run an AI on your own hardware, a different category of models is the one to look at — the one we track in the locally runnable section of our ranking, where fitting on a single machine is precisely the entry criterion.
”With no Nvidia chip at all”: what Z.ai said, and what it did not
This is the most repeated claim, and the most damaged in transmission.
What Z.ai actually claimed: the week-long preview and the inference that followed ran entirely on AI chips designed and manufactured in China, with a fleet of more than 100,000 units, at a cost per token comparable to Nvidia GPUs. Chinese manufacturer Cambricon announced immediate “Day 0” compatibility.
What Z.ai did not say: it named no manufacturer and no chip model in its announcement. It did not claim GLM-5.3-Flash had been trained on that hardware — the statement is about serving, not training. And it published no power consumption, no exact throughput, no utilisation rate. None of these results has been independently audited.
The difference is not a detail. Serving a model and training it are incomparable exercises: inference tolerates failures and parallelises cleanly, while training a frontier model requires tens of thousands of chips synchronised for weeks without a fatal interruption. Succeeding at the first does not demonstrate you can do the second.
So where does the “100,000 Huawei Ascend chips for training, zero Nvidia” figure that circulates attached to this model come from? From GLM-5, published on 11 February 2026: around 745 billion parameters, 28,500 billion training tokens, on a fleet of Ascend 910B with the MindSpore framework. That is a real and documented feat — but it is another model, released six months earlier. Transposing it onto GLM-5.3-Flash means crediting a company with a performance it did not claim for this product.
There is an irony nobody points out: the self-hosting documentation for GLM-5.3-Flash, published by Z.ai itself, asks for Nvidia Hopper. The model that symbolises freedom from Nvidia deploys, for those who download it, on Nvidia.
The subject remains serious: Chinese hardware autonomy is one of the structuring questions of the decade, and serving a model of this level without Nvidia is a real milestone. But it should be told with the facts that exist, not the ones we would like.
The MIT licence, and its big sister which is not one
The model’s most attractive argument is its licence. GLM-5.3-Flash is released under unmodified MIT: free download, free modification, free retraining, free integration into a commercial product, with no royalty and no authorisation. For a state, an administration or a company concerned about independence, that really is the least constraining form of dependence that exists.
That leaves the flagship. Its story is more interesting still, and it has barely been told in French.
GLM-5.3 did not come out on 28 August: it came out on the 14th, six days before Ox Alpha appeared. What was published on the 28th was its weights — and Z.ai had warned at launch that it would hold them back for about two weeks, to finish the safety evaluation and the hardening of the model.
The model itself is an industrial curiosity. It runs on exactly the same base as GLM-5.2, released in June: same architecture, same parameter count — some 750 billion, with public counts oscillating between 743 and 753 depending on what is included — and no retraining. All the gains come from extended post-training: more task environments, more environment types, for longer. At that level of performance, then, you can still progress a great deal without touching the base model. That is a notable fact, and it disappears completely behind the version number.
And that is where the reason for the delay becomes interesting. Z.ai explains that it added vulnerability discovery data hoping to improve reasoning on an isolated bug. What it got, it calls unplanned itself: the capability kept growing as training scaled up, and the model began reasoning across several exploitation steps, forming coherent plans for complete attack chains rather than simple bug hunting. On CyberGym it goes from 77.2% to 84.5%; on ExploitBench, from 24.4% to 54.4% — more than double in one iteration.
To sum up the sequence: the vendor delays publishing the weights because its model got better at offensive security than it had expected, then publishes them, downloadable by anyone. This is not an accusation: the same capability serves to find flaws before attackers do, and nobody asks a Chinese lab to be more cautious than its American competitors. But it is a subject we have already crossed on the side of AI-driven ransomware, and it is not settled by a contractual clause.
Let us look at that clause. The GLM-5.3 weights do not carry an MIT licence but an in-house one, whose only substantive restriction deserves reading: if the licensee or one of its subsidiaries operates a model-as-a-service business and aggregate revenue exceeds ten billion dollars over twelve consecutive months, it must pass Z.AI’s security review before any commercial use. Products that merely integrate model capabilities into a feature are not covered; nor is relaying requests to a third party’s hosted infrastructure.
The threshold therefore targets, in practice, only a handful of global players. But the clause is remarkable for another reason: the licence specifies that Z.AI alone determines the scope and method of that review, publishing neither criteria, nor deadline, nor route of appeal.
Translated: a very large European cloud provider wanting to build a commercial offering on Z.ai’s best model would have to obtain authorisation from a Chinese company — under a procedure whose rules are not published. That is exactly the kind of dependence the “open weights, therefore sovereignty” argument is supposed to remove.
Note above all what this licence does not do. It governs weights whose publication had just been delayed because of an unexpected offensive capability — and it says not a word about it. Its only barrier is a revenue threshold. And a ten-billion-dollar threshold filters hyperscalers, not attackers.
The commercial message is plain: the model that creates the buzz is MIT, the model that creates the revenue is not.
What our own ranking says — and the two ghosts we found in it
Ox Alpha never appeared in any ranking, and logically so: a stealth alias is attached to no vendor, and therefore to no evaluation source. Once published under its real name, the model entered normally. Here are the ranks recorded in our AI ranking at the update of 30 August 2026.
| Category | GLM-5.3-Flash | GLM-5.3 (flagship) |
|---|---|---|
| General | 106th | 70th |
| Code | 118th | 61st |
| Cost-performance | 84th | 76th |
| Chat | 139th | 97th |
Those ranks look modest against the Artificial Analysis index, and it is better to explain why than to hide them: our ranking aggregates around twenty public sources, and the same model family appears there under several variants depending on the evaluation protocol. It is a photograph of what the evaluators’ consensus publishes, not a verdict — and in the categories concerned, several hundred models are ranked.
Two honest clarifications while we are here. GLM-5.3-Flash does not appear in our “open source” category: that one lists, in practice, the models you can run on a local machine, with a size limit, and a 320-billion-parameter model falls out mechanically — its licence is not at issue. And in the “to watch” category, where it sits in second place, the rank does not follow the score: it is not a performance ranking.
But the most interesting thing we found while looking for something else.
In the code category of our own ranking there are still two models named Quasar Alpha (51st) and Optimus Alpha (57th). Their vendor is listed as “unknown”. These are two stealth models that circulated on OpenRouter in April 2025, before OpenRouter revealed they were early versions of GPT-4.1. The endpoints were withdrawn long ago. The rows are still there, with their unknown vendor.
In other words: more than a year after they were unmasked, two OpenAI models are wandering through the rankings under an anonymous identity, including ours. We will fix it. But it says something about the real traceability of this ecosystem.
Stealth is not an accident, it has become a method
Let us sum up what a lab gets by launching a model without saying its name.
It gets thousands of hours of free evaluation, in real conditions, on real code, in real projects, by motivated users — a volume of feedback no internal testing protocol can match. It gets it without exposing its brand: had the model disappointed, Ox Alpha would have vanished without a trace, and nobody would ever have known Z.ai had failed. And at the reveal it gets a story the world’s press picks up for free — a solved mystery is worth more than a launch release. The stock rise of 27 August is the direct measure of it.
And what does the user give? Their prompts.
This is not illegitimate, it is not concealed, and OpenRouter has formalised the practice in a dedicated programme. The Quasar Alpha precedent shows, moreover, that there is nothing specifically Chinese about the exercise: it was OpenAI that popularised it. But the thing has to be called by its name. A stealth preview is an exchange: free intelligence against usage data, under a contract the user cannot fully read, with a counterparty they cannot identify.
Presenting that exchange as a gift is the only part of the story that really causes a problem. And that is exactly what the video that started the wave did, with its line about “zero data retention”.
What this changes in practice
If you develop. The model is excellent and its value for money is hard to match today. Nothing forbids using it — it is now published, named, under an MIT licence, with downloadable weights. What should change is the reflex: before sending anything to an endpoint marked “stealth”, “preview” or “alpha”, open the data policy and the document it refers to. Both. That is three minutes of reading, and it is the only moment when you still have a choice.
If you run a company. The question is not “should we trust Chinese models?”, which is poor framing. The question is: who is named in your record of processing activities? A model under MIT licence hosted on your own servers, or at a European host you have contracted with, poses none of the problems described in this article. The same model queried through an anonymous endpoint poses all of them. It is not the model that creates the risk, it is the route by which you reach it.
If you simply follow AI news. Remember the distinction that derailed the video behind this article, and that none of the French-language pick-ups we read picked up on: “not used for training” does not mean “not kept”. You will meet it dozens of times, in the terms of service of products you use every day. It is probably the most profitable sentence to know how to read in the whole of 2026.
What to take away
Z.ai hid nothing illegal, lied about nothing, and ended up publishing everything: the name, the weights, the licence, the measurements. OpenRouter did not conceal its data policy — it was online, readable, in the expected place. The model really is remarkable, and its price is going to weigh on the whole market.
What failed was the reading. A twenty-word sentence, written in black and white, was turned into its opposite by a nineteen-second video, then picked up, then shared, until it became the information thousands of people believed they had at the moment they plugged their code into an unknown server.
There is a wider lesson. The AI ecosystem has invented a new object: the model that is simultaneously free, excellent and unidentified. The three qualities arrive together and that is no coincidence — it is the free price and the quality that buy the right to anonymity. As long as that triplet works, it will happen again — and it is one of the dynamics we track in our projection on what awaits AI in one, two and five years. Ox Alpha is neither the first nor the last; it is simply the one where the gap between what was written and what was understood was the most spectacular.
The good news is that the defence requires no technical skill. It requires opening the page and reading it to the end.
The figure to keep
Six. Six days between Ox Alpha’s appearance on OpenRouter and the moment we learned who owned the servers. Six days during which the model was free, with no serious usage limit, and plugged by thousands of developers into real projects. Six days of prompts retained by an entity nobody could name, under a contractual regime that two public documents described in two incompatible ways.
In the end, the provider was honourable, the model was good, and nothing untoward has been reported. That is luck, not a guarantee. The next time an anonymous and brilliant model appears — and it will — the only thing that will have changed is the number of people who thought to read the page before sending it their work.