It was supposed to be a quiet week, mid-July, while the West thinks about its holidays. Moonshot AI, a Chinese startup founded in 2023, decided otherwise: on 16 July 2026, it put Kimi K3 online, a 2,800-billion-parameter language model it presents as a head-on rival to Claude and GPT-5.6. Within hours, American, Taiwanese and Japanese technology and semiconductor stocks fell. A video circulating on social media sums up the mood: this model “floors every American digital giant”. Le Recul has checked what, in that announcement, genuinely holds up — and what is still a sales argument.

What Moonshot AI has just released

Kimi K3 is a mixture-of-experts (MoE) model: out of 896 internal “experts”, only 16 are activated for each computation, an architecture aiming to combine the power of a vast model with the compute cost of a far smaller one. Moonshot also claims an in-house hybrid linear attention named Kimi Delta Attention (KDA), a context window of 1 million tokens, and native vision support. The announcement coincides, as the source video explains, with the opening of the World AI Conference (WAIC) 2026 in Shanghai on 17 July. The timing is no accident: the day before, Beijing had also announced the creation of WAICO, an intergovernmental organisation for global AI governance bringing together 29 founding countries — a double demonstration, technical and diplomatic, that China now intends to weigh on the artificial intelligence stage, not merely follow it.

Kimi K3’s API has been available since 16 July 2026 (api.moonshot.ai, model kimi-k3), as well as on the Kimi app. On the other hand — and this is a point the video passes over quickly — the model’s weights, those that would allow it to be downloaded and installed on your own infrastructure, are not yet published: Moonshot promises to make them available on 27 July 2026, under a “modified MIT” licence. Until then, Kimi K3 is neither open source nor self-hostable, whatever part of the press may already imply.

The benchmarks: what is verified, what is inflated

The source video claims Kimi K3 “beats Claude Opus 4.8 on every benchmark” and ranks in the “top 3 on absolutely every benchmark”. The reality, as measured by third-party evaluators, is more nuanced. On Artificial Analysis’s Intelligence Index, an independent reference aggregating dozens of tests, Kimi K3 scores 57.1, which ranks it 4th out of 189 models evaluated — behind GPT-5.6 Sol and Claude Fable 5. On a specific programming benchmark, FrontierSWE, K3 scores 81.2, against 86.6 for Fable 5: on that precise test it does not “beat” it, it trails it.

Where the promise does check out is on one clearly identified leaderboard: LMArena, a blind comparison site judged by real users (the one the video calls “arena.ai”). Kimi K3 holds 1st place in the “Code – WebDev” category, a jump of 17 places on the previous generation (Kimi K2.6, then 18th). On its own set of 35 internal evaluations — therefore unverified by a third party — Moonshot claims around seven clear wins, including a score of 42.0 on the “SWE Marathon” test against 40.0 for Opus 4.8.

One figure deserves highlighting, because it directly tempers the breakthrough narrative: Kimi K3’s accuracy improves on the previous generation, rising from 33% to 46% correct answers on one test set — but its hallucination rate rises almost as much, from 39% to 51%. The model is right more often, but it also invents more often — a trade-off rarely highlighted in official announcements.

The price: cheaper per token, not necessarily cheaper per task

On paper, Kimi K3 is aggressive: 3 dollars per million input tokens, 15 dollars output (0.30 dollars for cached tokens) — against 5 / 30 dollars for GPT-5.6 Sol. It is also, independent developer Simon Willison points out, around triple the rate of its own predecessor K2.6 (0.95 / 4 dollars), and the highest rate ever charged by a Chinese lab for an open model. That per-token price matches, almost to the dollar, Claude Sonnet 5’s: Kimi K3 is therefore not cheap in absolute terms, only cheaper than the most expensive models on the market.

The argument that K3 would be “cheaper per task” thanks to better token efficiency does not entirely survive examination. On the Intelligence Index evaluation, Artificial Analysis measured consumption of around 130 million output tokens, more than double the median (63 million) of comparable reasoning models. Simon Willison, for his part, measured that a simple request could consume 13,241 reasoning tokens for only 3,417 tokens of useful answer — a ratio that pushes the bill for an innocuous prompt to around 25 cents, partly because only one reasoning effort level is available for now, unlike most recent competing models. The tokenizer itself seems poorly optimised: the same short prompt is counted as 95 tokens by Kimi K3, against 10 to 30 tokens at competing models for equivalent text — the sign of a hidden system prompt of around 85 tokens. Another detail worth noting for use in China: Moonshot applies separate pricing on its domestic market (around 2 ¥ for cached input, 20 ¥ for standard input, 100 ¥ output, per million tokens), a price difference between home market and international customers that is common among Chinese labs, but rarely highlighted in comparisons.

Market panic, then perspective

The comparison with DeepSeek (January 2025) imposed itself immediately on the markets: on the Kimi K3 announcement, technology and semiconductor stocks fell around the world on Friday 17 July — the Taiwan exchange lost more than 6%, Tokyo nearly 4%, the Nasdaq 1.5% (its worst session of the week), and the VanEck Semiconductor sector ETF (SMH) fell below its support moving average for the first time since April, more than 20% below its late-June record. Nvidia and Micron were particularly hit; figures such as David Sacks (the Trump administration’s AI adviser) and investor Bill Ackman publicly sounded the alarm.

Two qualifications are needed. First, most of the losses were recovered within the same day: most indices regained a good share of the lost ground by mid-session. Second, and above all, the DeepSeek analogy has a structural limit: the January 2025 DeepSeek shock (around 590 billion dollars of Nvidia market capitalisation wiped out in one session) rested on the idea that a frontier model could be trained at very low cost. Kimi K3, by contrast, is a premium model, priced in line with high-end American systems — which precisely weakens the argument that demand for compute, and therefore for chips, is collapsing. Le Recul had already followed that scenario during DeepSeek V4-Pro’s price cut: this time, the mechanism is not the same.

The handling of the story also diverges depending on where you look from. The Western press overwhelmingly took up the narrative of market panic and technological catch-up. Chinese state media (CGTN in particular) highlighted the global attention drawn by “Chinese open source AI” — a framing of technological and sovereign pride, with no mention of the same market worries. Both readings rest on the same figures, but keep different parts of them.

Who is Moonshot AI

Kimi K3 comes from Moonshot AI, a Beijing company founded in March 2023 by Yang Zhilin and two classmates from Tsinghua University. Yang Zhilin, a Tsinghua graduate with a doctorate from Carnegie Mellon (2019), is an alumnus of Google Brain and Meta AI, and co-author of the Transformer-XL and XLNet architectures — well-known references in language processing research, which sets his background apart from many less academically established AI startup founders.

The company’s financial trajectory is dizzying: valued at around 4.3 billion dollars in late 2025, it reached 10 billion in early 2026 after a 700 million raise, then 20 billion dollars in May 2026 following a 2 billion round led by Meituan, with Alibaba, Tencent, China Mobile and HongShan China among the investors — that is around 3.9 billion dollars raised in six months. Its annualised revenue (ARR) is said to have gone from more than 100 million dollars in March 2026 to more than 200 million by April, a progression to be taken with caution given how much that kind of figure measures, in this sector, an instantaneous billing rate rather than real profit. The company is also reportedly restructuring ahead of a possible Hong Kong flotation.

”Open weight” in ten days: what that really means

The source video is right on one precise point: the weights are not yet public, and the announced wait (around ten days) does match the timetable communicated by Moonshot, which targets 27 July 2026. But “open weight” does not mean “accessible to everyone”. Running a 2,800-billion-parameter model at home requires, according to the estimates reported, between 500 GB and 1 TB of memory and at least eight Nvidia H100 or H200 GPUs — a hardware bill in the order of 50,000 to 150,000 dollars. For an individual or a small organisation, the only realistic option therefore remains the API hosted by Moonshot, or by a reseller.

That point connects to an issue Le Recul has already documented around DeepSeek: using the hosted version of a Chinese model means passing through servers subject to the Chinese national intelligence law (2017), which can compel any organisation operating in China to cooperate with state services, regardless of the commitments displayed in a privacy policy. For personal, professional or sensitive data, the recommended caution remains the same as with previous Chinese models: keep hosted use to tasks with nothing at stake, and wait for a possible self-hostable version in Europe — itself then subject to the GDPR and the AI Act like any other infrastructure. Our open source model ranking and our AI coding model ranking will be updated as soon as the weights and independent evaluations are available.

What to take away

On 16 July 2026, Moonshot AI put Kimi K3 online, a 2,800-billion-parameter model (16 active experts out of 896), with a context of 1 million tokens.

On Artificial Analysis’s independent Intelligence Index, K3 ranks 4th out of 189 models, behind GPT-5.6 Sol and Claude Fable 5 — not “first on every benchmark” as the announcement says, even though it does hold 1st place on one specific leaderboard (LMArena, web code category).

Its accuracy improves on the previous generation, but so does its hallucination rate (39% → 51%).

The open weights, promised for 27 July 2026, will change nothing for most users: self-hosting such a model costs 50,000 to 150,000 dollars in hardware.

The announcement sent global technology stocks down before a partial rebound during the day — a “DeepSeek 2.0” scenario that the price does not really confirm: K3 is a premium model, not a knock-down one.

The figure to remember

51%.

That is the hallucination rate measured on Kimi K3 according to the evaluations reported by several independent technical analysts — up from its predecessor K2.6’s 39%, almost at the same pace as its accuracy gains (33% → 46%). A useful reminder: a model that is right more often can also, in the same movement, be wrong more often.