Two models released five weeks apart, and two opposite commercial strategies. On one side Claude Fable 5, published by Anthropic on 9 June 2026: proprietary, API-only, the most expensive on the consumer market — and currently top of the main independent intelligence index. On the other Kimi K3, published by Moonshot AI on 16 July 2026: 2.8 trillion parameters, the largest model ever released in open weights, downloadable free of charge since 27 July, and billed three times less on its official API.
The question is not which is “best” in the abstract — both are in the world’s leading pack. It is what you actually give up by paying three times less. Here is the answer, criterion by criterion.
Test results: the verdict in brief
Across our 7 criteria, Claude Fable 5 wins 4 to 3. It takes general reasoning, software-engineering reliability, speed and — the most important point in this comparison — factual soundness. Kimi K3 takes web interface coding, where it is first in the world, cost per task actually solved, where its lead is enormous, and deployment on your own infrastructure, where at this level of power it simply has no competitor.
| Criterion | Claude Fable 5 | Kimi K3 | Edge |
|---|---|---|---|
| General reasoning | Artificial Analysis index: 59.86 — first in the world | Artificial Analysis index: 57.11 — fourth | Fable 5 |
| Web interface coding | Arena.ai Frontend Code: 1,631 Elo (second) | Arena.ai Frontend Code: 1,679 Elo (first, in 6 of 7 domains) | Kimi K3 |
| Software engineering (reliability) | DeepSWE: 69.9% on the first attempt; more consistent across 4 attempts out of 4 | DeepSWE: 68.5% on the first attempt; better if 2 to 4 attempts are allowed | Fable 5 |
| Cost per task solved | 5.3 tasks solved per $100; $10 / $50 per million tokens | 14.7 tasks solved per $100; $3 / $15 per million tokens | Kimi K3 |
| Generation speed | 70.9 tokens per second | 32.0 tokens per second | Fable 5 |
| Factual reliability | No problematic hallucination rate reported; automatic fallback to Opus 4.8 on cybersecurity | Hallucination rate measured at 51%; 36% failure on problems with a logical trap | Fable 5 |
| Deployment and licence | Proprietary, API only, no self-hosting possible | Apache 2.0 open weights since 27 July 2026, self-hosting allowed | Kimi K3 |
Le Recul's verdict. Claude Fable 5 wins this comparison 4-3, but the score misses the point: these two models are not separated by power, they are separated by trust. The raw reasoning gap between them is 2.75 points out of 100 — negligible for most uses. The reliability gap is not.
Kimi K3 is the best value for money on the market, by a wide margin: 2.8 times more tasks solved per dollar spent, first place in the world on interface coding, and the only self-hostable option at this level of power. If your work tolerates a human check or a second attempt, the saving is hard to turn down.
But if a mistake costs you dearly, take Fable 5. A model that is wrong half the time and still answers with confidence is not a "cheaper" model: it is a model that moves the cost from the API bill to your proofreading time. That cost appears on no price list.
Two architectures, two philosophies
The headline figures pit a model whose size is unknown against a model that advertises its own as a selling point. That asymmetry is telling in itself.
Claude Fable 5 is the first “Mythos-class” model Anthropic has made public, a step above Claude Opus 4.8. Anthropic discloses neither the parameter count nor the architecture. What the vendor does document: a context window of one million tokens, up to 128,000 output tokens per request, and a claimed ability to hold a work thread across several million tokens by relying on notes it keeps as the task goes. A rarely mentioned design detail: on cybersecurity requests, Fable 5 switches automatically to Claude Opus 4.8, a less capable model. Anthropic states that this fallback fires in fewer than 5% of sessions.
Kimi K3 plays the opposite hand: everything is documented. 2.8 trillion parameters in total, but an extremely sparse Mixture-of-Experts architecture — the model has 896 experts and activates only 16 per token, which is 1.8% of its total capacity at each generation step. That is the model’s economic key: the size impresses, but the compute actually consumed stays under control. Moonshot adds three in-house pieces: Stable LatentMoE for routing, Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), designed to preserve the flow of information across long sequences and deep models. A one-million-token context as well, and native multimodal input (text and image), where Fable 5 is above all a text model.
Worth noting: Moonshot does not publish the exact number of active parameters. The 1.8% ratio is not enough to deduce it, because shared layers and non-expert components also consume compute. Any figure quoted on this point is an estimate, not official data.
General reasoning: Fable ahead, but the gap is thin
On the Artificial Analysis intelligence index, an independent reference that aggregates several evaluations, Claude Fable 5 scores 59.86 and takes first place in the world. Kimi K3 scores 57.11, in fourth — behind GPT-5.6 Sol max (58.89) and GPT-5.6 Sol xhigh (57.65).
2.75 points of difference out of 100. For scale: the median for comparable models sits around 32. Both are therefore at the very top, and the gap between them is far smaller than the gap separating them from the rest of the market.
Where Fable genuinely pulls away is on reasoning-dense tasks:
- SWE-Bench Pro (agentic coding): 80.3%, the best score ever published on this evaluation, against 69.2% for Claude Opus 4.8.
- Hebbia Finance Benchmark (senior-level financial reasoning): first of all models tested, with double-digit gains on document reading and on interpreting tables and charts.
Kimi K3, for its part, posts 89.2% on MMLU (general knowledge) and 94.2% on MGSM (maths translated into ten languages). Moonshot also announces 93.5% on GPQA Diamond, 91.2% on BrowseComp and 88.3% on Terminal-Bench 2.1 — these last three scores are self-reported by the vendor and have not yet been reproduced by an independent third party. We mention them; we do not count them in the verdict.
To place these two models against the rest of the market, our general AI ranking is updated continuously.
Coding: Kimi’s only clear win, and it is spectacular
This is where the comparison flips — but not everywhere.
Web interfaces: Kimi is first in the world
On Arena.ai’s Frontend Code Arena, where developers blind-compare the output of two models, Kimi K3 took first place in the world with 1,679 Elo points, ahead of Claude Fable 5 (1,631), GPT-5.6 Sol (1,618) and GLM-5.2 (1,587).
The detail makes the result more impressive still: Kimi K3 is first in 6 of the 7 domains evaluated — brand and marketing, design from a reference, data and analysis, consumer product, simulations, content-creation tools. It only yields first place on video games, where Fable 5 is ahead. And it is a jump of 17 places compared with the previous generation, Kimi K2.6, which sat in 18th.
See our AI ranking for code to place these two models against the specialised tools, and our comparison of the best AI for coding for the detail.
Software engineering: Fable takes back the lead on consistency
Together AI ran the most serious test published on this point to date: 452 scored runs across 113 real feature requests, drawn from live open source repositories, each attempted four times and validated by a hidden test suite.
The result:
- On the first attempt: Fable 5 succeeds 69.9% of the time, Kimi K3 68.5%. Fable’s edge, narrowly.
- Across 4 attempts out of 4: Fable solves more tasks consistently. It is more deterministic.
- If 2 to 4 attempts are allowed: Kimi moves ahead. It misses more often on the first try, but gets there in the end.
- Cost per run: $4.65 for Kimi against $13.41 for Fable at maximum effort.
The honest reading: Fable is the first-draft model, Kimi is the wide-net model. If your pipeline tolerates an automatic retry, Kimi wins. If every call has to land, Fable wins.
A real-world coding test published by The New Stack sums the equation up differently: same results, a third of the price, four times slower.
The real cost: the widest gap in this whole comparison
On the headline price, Kimi K3 is 3.3 times cheaper on output:
<tr style="background:#f8fafc;">
<td style="padding:11px; border:1px solid #e5e7eb;"><strong>Kimi K3</strong></td>
<td style="padding:11px; border:1px solid #e5e7eb;">$3</td>
<td style="padding:11px; border:1px solid #e5e7eb;">$15</td>
</tr>
| Model | Input per million tokens |
Output per million tokens |
|---|---|---|
| Claude Fable 5 | $10 | $50 |
Kimi also offers a cached input price of $0.30, a 90% discount on reused portions of context — useful for agents that reload the same corpus on every call.
But the headline price does not tell the whole story. The measure that counts is the cost per task actually solved, and that is where the gap becomes hard to ignore:
- Claude Fable 5: 5.3 tasks solved per $100
- Kimi K3: 14.7 tasks solved per $100
Kimi produces 2.8 times more useful work per dollar spent.
Two caveats, however, and they matter:
- Kimi reasons systematically, and those reasoning tokens are billed at the output rate. A chatty reasoning trace can cost more than the visible answer. The real bill therefore often exceeds an estimate based on the headline price.
- Code produced by Kimi needs manual correction more often, which shifts part of the cost onto human time — a cost that appears on no API invoice.
Our cost / performance ranking tracks this trade-off across the whole market.
Speed: Fable is twice as fast
70.9 tokens per second for Fable 5, 32.0 for Kimi K3. Fable is 2.2 times faster.
What that changes in practice:
- Live chat, customer support, autocompletion: the gap is visible to the naked eye. Fable’s advantage, no argument.
- Document analysis, batch processing, autonomous agents: latency barely matters, only total cost does. Kimi’s handicap becomes anecdotal.
It is a criterion often forgotten in model comparisons, and yet it alone is enough to rule out a choice depending on the use.
Reliability: the criterion that overturns the price argument
This is the most important point in this comparison, and the least covered in the reporting around Kimi K3’s launch.
Kimi K3’s measured hallucination rate is 51%, against 39% for the previous generation, Kimi K2.6. The model regresses on this point while progressing everywhere else. The technical explanation is coherent: K3 attempts more answers where K2.6 abstained. It therefore produces more right answers — and more wrong ones.
The real problem is not the rate, it is the behaviour when it is wrong: when Kimi K3 does not know, it answers with confidence more than half the time, instead of hedging or flagging its uncertainty. A flagged error can be caught; an asserted error spreads.
A second signal, sharper still: on problems containing a hidden invariant — an implicit logical trap that has to be spotted before answering — Kimi K3’s failure rate reaches 36%, against 8% for Claude Opus 4.8. That is an Anthropic model from a generation before Fable 5, already four times more solid on that specific exercise.
On the Fable 5 side, no problematic hallucination rate has been reported by third-party evaluators to date, and the architecture includes an explicit guardrail: the automatic fallback to Opus 4.8 on cybersecurity requests, triggered in fewer than 5% of sessions.
The practical translation: the 2.8× saving on the API bill is only real if you do not have to proofread twice as much output.
Languages: Fable’s advantage, with no dedicated score
Neither vendor publishes an evaluation specific to languages other than English. Here is what is established and what is not.
Claude Fable 5 is regularly described by independent comparisons as the stronger of the two on multilingual breadth. Anthropic does not publish per-language scores, but the model is trained on a broad multilingual corpus, and usage reports in French and other European languages point to nuanced reasoning.
Kimi K3 scores 94.2% on MGSM, a maths evaluation translated into ten languages. That is a good score, but it measures the ability to solve a problem stated in another language — not to write natural, idiomatic, stylistically sound prose in it. Moonshot mentions multilingual support without detailing it, and the training has clearly favoured English and Chinese.
A cautious conclusion: if a language other than English is central to your use — writing, customer relations, legal documents — Fable 5 remains the safer choice, pending dedicated evaluations. Our chat AI ranking follows this dimension.
Deployment: the one criterion where Fable has no argument
Claude Fable 5 is proprietary. Anthropic does not distribute the weights. Access goes through the API, without exception. Your data passes through the vendor’s infrastructure. There is no self-hosting option, at any price.
Kimi K3 is published in open weights under the Apache 2.0 licence since 27 July 2026, on Hugging Face. Apache 2.0 is one of the most permissive licences in existence: commercial use, modification and redistribution allowed, with no additional restrictive clause — unlike several so-called “open” licences on the market that impose user thresholds or usage restrictions.
What that allows in practice:
- No outbound data: the model runs on your servers, nothing leaves for the vendor.
- No vendor lock-in: the model keeps working even if the vendor shuts down, changes its pricing or restricts access.
- Zero marginal cost per use: beyond a certain volume, only the infrastructure counts.
And what it costs:
- Around 594 GB to download for the native MXFP4 version.
- Serious GPU infrastructure, out of reach of a workstation or a small server.
- The expertise to run it (vLLM or equivalent), supervise it and maintain it.
The break-even point is real but high. For the vast majority of projects, the Kimi API remains cheaper than self-hosting. Self-hosting becomes rational in two cases: very high volume, or a confidentiality constraint that forbids data leaving your infrastructure. Our ranking of open models lists the alternatives against these criteria, and our comparison of a free agent against a paid one applies the same reasoning to autonomous agents.
Who should choose what
Take Claude Fable 5 if:
- a mistake costs you dearly — finance, health, legal, production;
- you work mainly in a language other than English;
- your use is interactive and latency shows;
- you need the best possible first draft, with no retry;
- your budget absorbs $10 / $50 per million tokens.
Take Kimi K3 if:
- you build web interfaces — that is its speciality, and it is first in the world at it;
- your volume is high and your budget constrained;
- your pipeline tolerates a second attempt or an automatic check;
- you have to host the model yourself, for confidentiality or independence;
- you work in batches, where slowness costs nothing.
Both work if you are doing everyday text generation, summarising or simple question-answering, with no critical reliability requirement and no strong language constraint.
What this comparison says about the market
At the end of July 2026, the raw power gap between the best proprietary model and the best open model has fallen to 2.75 points out of 100. That is the real headline, more than the ranking itself.
The premium argument can no longer rest on power alone: it now rests on reliability, speed and consistency. That is precisely where Fable 5 keeps its lead — and precisely where Kimi K3 will have to improve to turn its pricing win into an outright win.
Our other comparisons apply the same method to other match-ups on the market.