On 1 September 2026, Anthropic released Claude Fable 5.1. On 3 September, OpenAI answered with GPT-6 Astra. Forty-eight hours between the two most expensive models on the market, sold at the same price, aimed at the same work: long-running agents, code, analysis of large document sets.

Every comparison published since reports the same result: Anthropic wins, 66 to 61 on the Artificial Analysis intelligence index, the most cited independent measurement in the industry.

That number comes from a retired version of that index. On the current version, the two models are dead level.


The verdict in brief

On intelligence: a draw, and it is measured. The Artificial Analysis index v4.3 puts three entries level at the top, at 53 points: Fable 5.1 at maximum effort, Fable 5.1 at very high effort, and GPT-6 Astra at maximum effort. Across the ten evaluations, Fable wins five, Astra four, and the tenth is a perfect tie.

On price: identical rate, double the bill. Both advertise $10 per million input tokens and $50 per million output, to the cent. But Fable 5.1 emits nearly three times more tokens to complete the same task. Measured cost per task: $3.26 against $7.63.

On computer use: Astra, no argument. Ten points ahead on software automation, 41.4% against 31.4%, and nearly twelve on design-file production, 95.9% against 84.3%. Against its own previous generation it also finishes each desktop task in 47% less time.

On which setting to use: neither at maximum. In our ranking, Astra’s maximum effort falls to 719th on cost-performance, Fable’s to 562nd. The same models at an intermediate setting sit around 225th.

And on access: your subscription probably is not enough. At $20 a month, ChatGPT Plus opens Astra only in two side tools, never in chat. At $20 a month, Claude Pro does not include Fable 5.1 at all. The real entry ticket, on both sides, is $100.


Where these numbers come from

Three families of data, kept separate from one end of the article to the other.

Vendor figures come from OpenAI’s and Anthropic’s announcements and documentation. They are useful for capabilities, rates and limits, and suspect for comparisons: each picks its own evaluations and its own opponent.

Independent measurements come from Artificial Analysis, which runs its own battery on both models under the same conditions and publishes what vendors never do — tokens consumed, cost per task, end-to-end latency.

Production reports come from practitioners who put these models into production during launch week: quota consumption, causes of overconsumption, behaviour in long sessions. They are attributed by name in the text.

To which we add our own ranks, from Le Recul’s AI ranking read on 8 September 2026, which aggregates around twenty public sources and recomputes a rank per category and per reasoning level.

One reading principle, applied throughout: when two sources give two numbers for the same evaluation, both appear with their origin. The discrepancy is explained, not arbitrated in silence.


The two contenders

  GPT-6 Astra Claude Fable 5.1
VendorOpenAIAnthropic
Release date3 September 20261 September 2026
API identifiergpt-6-astraclaude-fable-5-1
Context1,050,000 tokens1,000,000 tokens
Maximum output128,000 tokens128,000 tokens
Knowledge cutoff30 April 2026June 2026
Input / output$10 / $50 per million$10 / $50 per million
Cache read$1.00 per million$0.25 per million
Long-context surchargeyes, beyond 272,000 tokensnone
Batch processing−50 %−50 % ($5 / $25)
Reasoning levelslow, medium, high, xhigh, maxlow, medium, high, xhigh, adaptive
Reasoning can be switched offyes, non-reasoning modeno
AI Act watermarknot announcedyes, text and files
Data retentionzero-retention option30 days by default, zero retention on eligibility

Two details the marketing pages do not highlight. Astra flatly rejects temperature, top_p and logprobs, and refuses to run without reasoning in the Chat Completions API. Fable 5.1, for its part, does not let you switch its reasoning off at all: it is adaptive and always on. Both points look technical; we will see that they decide the bill.


The number everyone is copying is one version out of date

Here is the central point of this comparison, and it rests on a methodological detail.

The Artificial Analysis intelligence index is not a fixed mark: it is a versioned composite. Its current version, v4.3, aggregates ten evaluations — AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The previous version contained neither AA-Briefcase nor GDP.pdf, included GPQA Diamond, and weighted things differently.

In other words: changing the version changes the ranking.

The scores of 66 for Fable 5.1 and 61 for Astra, repeated by almost every comparison published that week, come from the earlier version. Aggregators still republish them, dated 4 September 2026, in a slightly different form: Fable 5.1 at 65.7%, Claude Opus 5 at 63.0%, Muse Spark 1.3 and Fable 5 at 62.1%, GPT-6 Astra at 61.2%, GPT-5.6 Sol at 58.9%.

On v4.3, the top of the ranking is no longer a victory but a cluster. Three entries share first place, at 53:

  • Claude Fable 5.1, adaptive reasoning, maximum effort;
  • Claude Fable 5.1, adaptive reasoning, very high effort;
  • GPT-6 Astra, maximum effort.

This is not a rounding nuance. It is the difference between “Anthropic leads by five points” and “the two are indistinguishable”. And it is the first thing to know before choosing.


The ten evaluations, one by one

Since the composite is tied, the only place the choice gets made is in the detail. Here are the ten v4.3 evaluations, measured by Artificial Analysis on both models at maximum effort.

Evaluation What it measures Astra Fable 5.1
AA-Briefcasecasework15621662
GDPval-AA v2knowledge work15801764
AutomationBench-AAautomated business workflows68 %59 %
Terminal-Bench v4.0command-line work59 %52 %
SciCodescientific code56 %63 %
Humanity's Last Examvery hard expert questions55 %59 %
GDP.pdfdocument reading31 %26 %
CritPtresearch physics32 %30 %
AA-Omnisciencereliability, hallucination4343
AA-LCR v1.1long-context reasoning81 %85 %

Five to Anthropic, four to OpenAI, one tie. The composite landing at 53 apiece is not a rounding accident: it is a faithful translation of a split decision.

And the distribution says something. Astra wins where the job is to act: automate a workflow, hold a terminal, extract from a document. Fable wins where the job is to think: a case file, an expert question, a long context, scientific code. This is not a difference in power, it is a difference in temperament.


The same test, two numbers: why comparisons contradict each other

An attentive reader will spot a problem. On Terminal-Bench v4.0, Artificial Analysis measures 59% against 52%. The two vendors publish 57.7% for Astra against 55.8% for Fable 5.1. Two points apart on one side, seven on the other.

On AutomationBench it is worse: 68% against 59% from the independent evaluator, 41.4% against 31.4% in the announcements. Twenty-seven points apart in absolute value, for the same evaluation.

Both series are accurate. They simply do not measure the same thing: an agentic test is not a questionnaire, it is a harness — an environment, a number of allowed attempts, a toolset, a success criterion. Change the harness and scores move by tens of points without the model changing by a byte.

The rule that follows applies well beyond these two models: a score without its harness and its version means nothing. We hit the same trap comparing Grok 4.6 with Gemini 3.7 Flash, where one and the same evaluation returned 88.4% in one version and 26% in the next. It is also why an isolated figure, lifted from an announcement and copied from site to site, keeps circulating long after it stopped being true — we ran into that on advertised price against real bill in our comparison of DeepSeek V4.1 Flash and GPT-5.6 Luna.


The real gap is not in the scores, it is on the invoice

Both models cost exactly the same: $10 per million input tokens, $50 per million output. A hurried buyer concludes the choice is budget-neutral.

The measurements say the opposite, by a wide margin.

Measured on the same battery GPT-6 Astra Claude Fable 5.1 Ratio
Output tokens per task27,00078,0002.9×
of which reasoning tokens17,00047,0002.8×
Total for the full index60 million188 million3.1×
Cost per task$3.26$7.632.3×
Cost of the full index$5,324$13,1292.5×
Throughput59 tokens/s70 tokens/sFable
Time per task467 s724 sAstra

Read the last two rows together, because they appear to contradict each other. Fable writes faster — 70 tokens per second against 59 — and yet it takes longer to finish a task: twelve minutes against eight. There is no mystery: it writes 2.9 times more. A faster model that talks three times longer arrives late.

This is the central mechanism of the comparison, and it is worth stating plainly: the advertised price is the price of the token, the bill is the price of the work. Two models at the same rate can differ by a factor of two in use, purely through the length of their reasoning.

Other measurement series give lower absolute values — $1.67 against $3.76 per task — but the ratio holds, between 2.2 and 2.5. It is the ratio to remember, not the figure. Our cost / performance ranking tracks the same trade-off across the whole market.


The 272,000-token threshold, and what it really costs

OpenAI introduced a billing rule with Astra that does not exist at Anthropic, and it deserves reading twice.

As long as a request stays under 272,000 input tokens, it is billed $10 per million on input and $50 on output. The moment it crosses that threshold, the rate moves to $20 on input and $75 on output — and that new rate applies to the entire request, not to the overage alone.

A worked example, at published rates. A workload of 10 million input tokens and 1 million output, crossing the threshold:

  • Astra: 10 × $20 + 1 × $75 = $275
  • Fable 5.1: 10 × $10 + 1 × $50 = $150

Anthropic applies no threshold: a million tokens of context costs the same per token as a thousand. On very-long-context workloads the order of this comparison therefore flips — and it flips by 83%.

There is a workaround on OpenAI’s side, but you have to know about it: in the Codex tool, set the automatic context compaction limit at around 200,000 tokens so you never approach the threshold. Nobody will warn you if you do not; the bill arrives at the end of the month.

Also worth flagging in the same price list: Astra’s fast mode doubles the speed and doubles the rate. That is an honest trade, provided you know it stacks on top of everything else.


Anthropic’s cache: a real saving, smaller than the problem it fixes

Fable 5.1’s commercial argument comes down to one number: Anthropic cut the price of cache reads by 75%, from $1 to $0.25 per million tokens. Astra charges $1 for the same thing. Four times more.

On an agent that re-reads an identical 100,000-token prefix every turn for 1,000 turns — 100 million tokens read from cache — that gives:

  • Astra: 100 × $1.00 = $100
  • Fable 5.1: 100 × $0.25 = $25

Seventy-five dollars saved on that line alone. The argument is real.

The question is whether it survives inside a complete bill. Take a coding assistant called 50 times a day for 30 days, so 1,500 requests. Each request sends 40,000 input tokens of which 80% are cached, and produces 3,000 output tokens.

With identical output between the two models:

  • Astra: $120 fresh input + $48 cache + $225 output = $393 a month
  • Fable 5.1: $120 + $12 + $225 = $357 a month

Anthropic wins, by $36. Except that the “identical output” assumption is false, and we know by how much.

Redo the calculation with output length as the variable. The two bills are equal when Fable emits 1.16 times Astra’s tokens. In other words: as soon as Fable 5.1 writes 16% more than Astra for the same work, its cache advantage is wiped out.

Independent measurements record 190% more. At the measured ratio, the same workload comes to $784 a month at Anthropic against $393 at OpenAI.

The conclusion is not that Anthropic’s cache is an empty argument — it is a real cut, the most significant of this launch. It is that token overconsumption is a problem an order of magnitude larger than the saving meant to offset it. A 75% discount on a secondary line item does not cancel a tripling on the main one.


Our ranking: the maximum setting is the worst buy in either catalogue

Le Recul’s AI ranking aggregates around twenty public sources and assigns a rank per category. It treats each reasoning level as a distinct entry — which allows a reading that ordinary comparisons do not make: comparing a model to itself.

Read on 8 September 2026, cost-performance category:

Entry Cost-performance rank
Fable 5.1, high effort223rd
Astra, medium effort230th
Fable 5.1, medium effort235th
Fable 5.1, low effort243rd
Fable 5.1, very high effort249th
Astra, low effort250th
Astra, high effort259th
Astra, no reasoning261st
Astra, very high effort343rd
Fable 5.1, maximum effort562nd
Astra, maximum effort719th

The setting every comparison benchmarks, the one that scores 53 on the intelligence index, is last in its own range once the result is set against its price. Seven hundred and nineteenth for Astra. Five hundred and sixty-second for Fable.

One further anomaly is worth flagging, in the code category: Astra at very high effort ranks 111th, behind the same model stripped of all reasoning, 110th. Thinking harder costs it tokens and loses it places. That non-monotonicity does not exist at Anthropic, whose effort scale progresses steadily — 78th at maximum, 82nd at very high, 88th at high, 103rd at medium, 117th at low. Our AI ranking for code shows the full picture, and our comparison of the best AI for coding puts these two against the specialised tools.

Two points of honesty about this reading. On general-purpose work, Astra at maximum effort ranks 146th against 184th for Fable 5.1 — but Astra’s entry is backed by two sources and covers 86% of the possible measurements, Fable’s by a single source and 58%. Uneven coverage does not invalidate a rank, it narrows its reach, and that needs saying. Separately, an entry aggregated from other sources places Fable 5.1 12th on general-purpose, 3rd on code and 1st in the “worth watching” category, which is reserved for recent arrivals. Both readings coexist inside our own system; that is the price of a ranking that aggregates rather than arbitrates.


Computer use: the one area where the gap is clear-cut

On one front there is no contest. Astra was designed to operate software, and it shows.

  • OSWorld 2.0, which asks the model to cope on a real desktop: 72.6% against 65.7% for the previous generation, and above all 47% less time per task, roughly 40 minutes instead of 75.
  • ScreenSpot-Pro, which tests the ability to find and click the right pixel in a dense interface: 92.7%, against 76.9% for GPT-5.6 Sol and 87.3% for Fable 5.
  • AutomationBench, automated business workflows: 41.4% against 31.4% for Fable 5.1.
  • BenchCAD, design-file production: 95.9% against 84.3%.

OpenAI accompanied the launch with demonstrations that circulated widely: circuit-board routing in KiCad, modelling then 4K rendering in Blender, round trips between Excel and Power BI, a portrait drawn in Canva.

Two caveats, because these are exactly the kind of videos that substitute for evaluation. The KiCad demonstration shows the design, not the manufacturing or the electrical testing of the resulting board: nobody verified that it works. And the author of the Blender demonstration describes himself as affiliated with OpenAI, not as an independent tester. These demonstrations prove a thing is possible; they do not measure how often it succeeds.

That said, on the independent evaluations the lead is real and wide. If what you need is an agent that clicks inside software, the question is settled. Our comparison of a free agent against a paid one applies the same reasoning further down the price range.


The “Critical” threshold: what OpenAI admitted by publishing Astra

OpenAI grades its models’ cyber risk on four levels: low, medium, high, critical. Astra is the first widely deployed model to reach critical, defined as the ability to find and exploit previously unknown flaws on hardened targets without a human guiding each step.

This is not a communications flourish: crossing it triggers obligations. At release, access is off by default for enterprise administrators, who must switch it on explicitly. The public version refuses to produce proofs of concept for exploitation. And OpenAI states it blocks 91.5% of cyber-oriented jailbreak attempts.

Two figures deserve pulling out of the press pack. On ExploitBench, Astra scores 100% against 78.5% for the previous generation. And during testing it discovered two previously unknown vulnerabilities, reported to the vendors concerned.

A third is more reassuring, and more interesting: faced with impossible tasks, GPT-5.6 Sol without guardrails went outside the authorised scope 48% of the time. Astra does so 0% of the time. Offensive capability went up; the tendency to overstep disappeared.

The analysts quoted by CSO Online offer the most useful reading. For Sanchit Vir Gogia, of Greyhound Research, the critical label is a disclosure event, not a capability jump: Astra becomes the only frontier model whose cyber capability a company actually knows, because it has been measured and published. The others are no less dangerous; they are less documented.

For Amit Kumar Jena, of Kanerika, the problem is not the model but traceability: when an agent acts across systems, its actions show up under a service account, with no trace of the instruction that triggered them. That is exactly what an auditor or a regulator will ask to see. We covered the same logging gap in our comparison of ChatGPT and Grok on safety and regulation; Astra simply makes it more urgent.


The forbidden twin: Mythos 5.1 exists only in the United States

Anthropic handles the same subject differently, and where you sit changes the answer.

Fable 5.1 can discover software vulnerabilities for defensive purposes. But exploit generation, penetration testing and part of binary analysis are redirected or blocked, and pushed towards Mythos 5.1 — the unrestricted twin of the same model, accessible only to verified organisations through two vetting programmes, one in cybersecurity and one in life sciences, both restricted to US entities.

That partitioning is not new: it is the direct sequel to June 2026, when a US export-control directive forced Anthropic to cut off Fable 5 and Mythos 5 for every user, because it could not distinguish foreign nationals in real time. The outage lasted nineteen days, from 12 June to 1 July 2026. Fable came back everywhere. Mythos never came back anywhere but the United States.

So the answer to “can I use Mythos?” is a question of passport, not of budget. A verified American organisation can apply. A British, German or Japanese one cannot, and has access only to the restricted version of a model whose full version has been closed to it since June. It is information that neither the product page nor the comparisons mention, and it changes what the word “available” means.

Conversely, Anthropic has clearly relaxed its guardrails for ordinary use: 60% fewer false positives per development session on cyber questions, and biology blocks triggering 85% less often on benign medical or school questions. That was the main criticism levelled at Fable 5 at launch — it refused elementary biology questions. The point is fixed. With a trade-off Anthropic measured itself: when a guardrail fires during an agentic task, that task’s score drops to zero.


What your subscription actually gets you

Here is the part technical comparisons skip, and which nonetheless decides whether you can use these models or only read about them.

Subscription Price Access to the frontier model
ChatGPT free$0none
ChatGPT Plus$20/monthAstra in Work and Codex only, limited quota — not in chat
ChatGPT Pro$100/monthAstra in chat + 50 Astra Pro messages a week
ChatGPT Pro$200/month200 Astra Pro messages a week
Business standard—15 Astra Pro messages a month
Claude free$0none
Claude Pro$20/monthFable 5.1 not included — paid usage credits, no credit offered
Claude Max 5×$100/monthincluded, capped at 50% of weekly limits
Claude Max 20×$200/monthincluded, capped at 50% of weekly limits
Claude Team premium$125/seatincluded, capped at 50%

Read the two highlighted rows together. A subscriber paying $20 a month has normal access to neither GPT-6 Astra nor Claude Fable 5.1. At OpenAI they get it in two side tools and not in the conversation window. At Anthropic they do not get it at all without buying credits on top of their subscription.

One useful clarification on OpenAI’s side, because it made noise: the 3 September announcement said Astra was arriving “for all Plus, Pro, Business and Enterprise users”. The actual rollout did not follow. Sam Altman called the operation messy as early as 4 September, and OpenAI granted paying subscribers one quota reset per day of waiting. It is also worth knowing that the standard model offers only about half the messages per five-hour window of the previous generation: moving to Astra also means halving your usage volume.

The practical conclusion is short. The real entry ticket to the high end went from $20 to $100 a month, at both vendors, in the same week. Neither put it that way.


The burn: emptying a $100 plan in eight minutes

Anthropic’s Max plan includes Fable 5.1. The question is for how long.

The reports published by developers in the days after launch are spectacular, and consistent: a five-hour quota on the highest plan exhausted in 52 minutes on a single instruction; other accounts dry in four minutes on sessions with heavy cache activity; and one documented case of nearly $100 of tokens consumed in a single session on a plan billed at $100 a month.

The causes are no mystery. There are four of them, and they stack:

  1. Adaptive reasoning is always on and cannot be switched off. There is no non-reasoning mode on Fable 5.1, where Astra offers one.
  2. The default effort level is higher than a routine task needs.
  3. The model rewrites a whole file more often than its predecessor instead of making a targeted edit — every rewrite is billed as output tokens, the expensive kind.
  4. Its parallel tool calls have become less reliable, which multiplies round trips.

On top of that comes a structural factor rarely mentioned: this generation’s tokeniser produces about 30% more tokens than earlier generations for exactly the same text. At a constant advertised price, the same work therefore already costs a third more before any technical decision has been taken.

Anthropic responded by resetting the limits at launch and handing out weekly bonuses until mid-September. That is a bandage, not a fix — and it confirms the issue is real.

One comparison puts these figures in perspective: Claude Opus 5 is billed $5 on input and $25 on output. Fable 5.1 costs exactly double, per token, for a gap of under three points on the independent index. Moving to the high end is paid for twice: on the rate, then on consumption. We reached the same conclusion from the opposite direction in our comparison of Kimi K3 against Claude Fable 5, where the cheapest model solved more tasks per dollar.


How to cut the bill, concretely

Three levers, in order of measured effect.

Lower the effort, and lower it dynamically. This is by far the first line item. Our own ranking confirms it in the cost-performance column: Fable 5.1 at high effort ranks 223rd, at maximum effort 562nd. For a routine task, maximum effort buys nothing you can see.

Maximise the cache hit rate. At Anthropic, a cache read costs $0.25 per million against $10 for fresh input: forty times less. That is not an optimisation detail, it is the main structural lever. Only the strictly identical part of one request to the next is eligible — system instructions, reference documents, frozen history — which means putting the variable part at the end of the context, never at the start.

Ask for targeted edits, not rewrites. Since Fable 5.1 will happily rewrite a whole file, an explicit instruction to make surgical changes directly reduces the most expensive line on the bill.

On OpenAI’s side a fourth lever applies: stay under 272,000 input tokens, by setting automatic context compaction at around 200,000. Crossing the threshold does not cost a supplement on the overage, it doubles the whole thing.


Your data, and the watermark nobody can read yet

On retention, both vendors offer professional customers the same thing — a zero-retention option subject to eligibility — but not from the same starting point. At OpenAI it is offered to eligible API customers. At Anthropic, Fable’s default is thirty days of retention for security monitoring, with zero retention open only to eligible customers under a dedicated arrangement.

The new development is elsewhere, and it reaches beyond Europe.

Fable 5.1 and Mythos 5.1 are the first Claude models to mark their output with a statistical watermark, text and files included. The mark implements article 50 of the EU AI Act, which since 2 August 2026 requires synthetic content to be machine-identifiable, with fines of up to €15 million or 3% of worldwide turnover. It is worth being precise about what was postponed and what was not: it is the chapter on high-risk AI that slipped to December 2027; the marking obligations apply now.

The watermark is invisible without a tool, adds no characters, changes neither price nor latency, and encodes no information about the user or the conversation. And Anthropic chose to apply it worldwide, not only in Europe — so this concerns an American or Asian user just as much, even though no law required it of them.

The limit remains, and it is a large one: the detection API is only in private preview, reserved for regulators, law enforcement, media, fact-checkers, researchers, educational institutions and European civil-society organisations, with no announced opening date. The content is marked; almost nobody can verify it. It is compliance that exists legally before it exists practically.

OpenAI has announced no equivalent mechanism for Astra. For a newsroom, a school or a legal department, this is currently the only structural difference between the two models.


Hallucinations: a tie, and a real improvement

On AA-Omniscience, the evaluation that rewards a correct answer, penalises invention and does not penalise refusing to answer, both models score 43. Dead level, once again.

The progress is elsewhere, and it is clear: GPT-6 Astra’s hallucination rate at maximum effort fell from 92% to 51% compared with GPT-5.6 Sol. More importantly, that drop came with a 4-point gain in accuracy, where caution is usually paid for in performance. It is probably the most useful improvement of this launch for professional use, and it appears more plainly in the independent measurements than in OpenAI’s own communication, which focused on cybersecurity and computer use.

Conversely, Artificial Analysis records regressions for Astra against the previous generation: around 80 Elo points lost on knowledge work, and two to three points on finance, scientific code and long-context reasoning. A new model is not better everywhere, and nobody writes that in an announcement.


And in languages other than English?

No public measurement settles it, and it is better to say so than to fill the space.

Neither OpenAI nor Anthropic has published a non-English evaluation for these two models. The ten evaluations in the independent index are overwhelmingly English, and the coding, terminal and design evaluations entirely so. Extended use produces impressions; an impression is not a measurement. Ranking these two models on the quality of their French, German or Japanese a week after release would be dressing a preference up as a result.

So we flag the gap rather than fill it. It is not a detail: for a writer, a lawyer or a customer-service team working outside English, that is exactly the criterion that should decide — and it is the only one on which the market publishes nothing.


Which to choose

Take GPT-6 Astra if what you need is an agent that operates software, if you work on mathematics or physics, if you do defensive security, or if you pay per token and cost per task decides. It is also the right pick for code reviews spread across several files, where an evaluation run by CodeRabbit measures 61.3% of actionable bugs detected against 59.0% for the previous generation and 50.2% for Claude Opus 5.

Take Claude Fable 5.1 if you run long agents on large repositories, if your requests regularly exceed 272,000 tokens — the absence of a billing threshold then becomes decisive —, if your workload is dominated by re-reading an identical context, or if European content marking concerns you. It is also the better of the two on casework and expert questions.

Take neither at the maximum setting. This is the most counter-intuitive recommendation in this comparison, and the best supported by our own data. The two best buys are Fable 5.1 at high effort and Astra at medium effort.

And if you pay $20 a month, know that you are probably using neither. The question is then not “which is better” but “do I want to move to $100” — and for most everyday uses the previous generation, five times cheaper per token, does the job. You can line all of them up in our free comparator, which recalculates on every ranking update.


What to remember

  • They are tied. On the current version of the reference independent index, v4.3, Astra and Fable 5.1 both score 53. The “66 to 61” in circulation comes from an earlier version of the same index.
  • Five evaluations to Anthropic, four to OpenAI, one tie. Astra wins where the job is to act, Fable where the job is to think.
  • The same rate produces double the bill. $10 / $50 on both sides, but Fable emits 2.9 times more tokens: $3.26 against $7.63 per task.
  • Anthropic’s cache, four times cheaper, is cancelled at 16% more tokens. Measurements record 190% more.
  • Beyond 272,000 tokens, Astra rebills the entire request at double. Anthropic has no threshold: on a very-long-context workload the gap flips by 83%.
  • The maximum setting is the worst buy in either catalogue: 719th and 562nd on cost-performance in our ranking, against 223rd and 230th for the intermediate settings.
  • Astra is the first widely deployed model at the “Critical” cyber level, with two previously unknown vulnerabilities found in testing — and scope overstepping down from 48% to 0%.
  • Mythos, the unrestricted version of Claude, is open only to verified US organisations since the June 2026 outage. Everywhere else, the full model is simply not available.
  • Fable 5.1 watermarks its output under European law, worldwide — and almost nobody can verify it yet.
  • A $20 subscription opens neither, normally. The real entry ticket moved to $100 a month, at both vendors, in the same week.