On 28 May 2026, Anthropic launched Claude Opus 4.8.
At first glance this looks like an ordinary update: better reasoning, better code, better benchmarks.
But behind the benchmarks sits a far more important shift.
Anthropic is no longer only selling a more capable AI.
The company is starting to sell an AI that is:
- more autonomous;
- able to work for longer;
- able to coordinate hundreds of sub-agents;
- and above all presented as more honest.
An unusual pitch in an industry that normally communicates about performance rather than mistakes.
An AI that admits its limits more often
Anthropic’s main message is surprising.
The company says it worked specifically on the model’s ability to:
- recognise its uncertainties;
- avoid jumping to conclusions;
- flag more often when information is missing;
- spot its own errors more easily.
According to Anthropic, Opus 4.8 is roughly:
4 times less likely to let certain errors in its own code pass unflagged.
That figure should be read in the context of Anthropic’s internal tests.
But it points at a reality that is often forgotten:
AI is progressing fast, yet hallucinations and reasoning errors remain a major problem.
And the more autonomy AI gains, the more sensitive that problem becomes.
Dynamic Workflows: the real novelty
The most important feature is probably not the model itself.
Anthropic is introducing Dynamic Workflows, currently in preview.
The principle:
Claude can now:
- break down a complex task;
- create several specialised sub-agents;
- run those sub-agents in parallel;
- compare their results;
- merge the information before answering.
Anthropic now talks about:
hundreds of parallel sub-agents in a single session.
Two years ago, the main subject was the prompt.
Today, the subject is becoming the orchestration of swarms of agents working together.
That is not an isolated Anthropic move either: measured across the whole ecosystem, the tools agents use have shifted from reading to acting in barely sixteen months.
Claude works for longer
Another important change:
Anthropic is pushing Claude towards longer and longer tasks.
The company says Claude can now handle:
- complex analyses;
- large software projects;
- extended workflows;
- massive code migrations.
Anthropic even mentions:
hundreds of thousands of lines of code handled within a single project.
So the subject is no longer simply:
“Claude answers better.”
The subject becomes:
“Claude works for longer before a human steps in.”
And that sentence describes a job change more than a product change. Some developers were already saying it out loud the day before this launch: they spend more time supervising agents than writing code.

The giant context remains a major advantage
Claude keeps one of its main strengths.
Maximum context: 1 million tokens
For comparison:
- Claude Opus 4.6 had already opened the way to a 1M context;
- Claude Opus 4.7 kept it;
- Claude Opus 4.8 keeps it as a baseline.
So the model can analyse:
- several hundred pages;
- entire code bases;
- thousands of files;
- extremely long conversations.
Anthropic also maintains:
Maximum output: 128,000 tokens
Which remains one of the largest generation capacities on the market today.
Consumption that becomes more flexible
Anthropic is not changing Opus’s standard price compared with Opus 4.7.
Standard API pricing
- Input: $5 per million tokens
- Output: $25 per million tokens
Fast Mode pricing
- Input: $10 per million tokens
- Output: $50 per million tokens
This point matters: Opus 4.8 is not sold at a higher price than its predecessor in standard use.
What can vary more, however, is actual consumption, depending on the effort level chosen.
The user can ask for:
- faster thinking;
- deeper thinking;
- longer execution;
- or a mode better suited to hard tasks.
The longer Claude thinks:
- the more tokens it consumes;
- the more budget it uses;
- but the more robust the result can be.
Anthropic is gradually turning reasoning into an adjustable resource.
Fast Mode: faster, but not free
Anthropic is also announcing an accelerated mode.
According to the company, Fast Mode can reach:
- up to 2.5 times faster;
- at a specific rate of $10 / $50 per million tokens.
The goal is clear:
make Opus more usable in situations where speed matters as much as quality.
But it is also a reminder of something often forgotten:
advanced AI does not only cost by model.
It also costs by duration, by effort, by context, by number of agents and by token volume.
Claude Desktop becomes strategic
This release also confirms an underlying trend.
Claude is no longer simply a chatbot.
With:
- Claude Desktop;
- Claude Code;
- MCP;
- Dynamic Workflows;
Anthropic is gradually building a platform capable of interacting with:
- files;
- software;
- browsers;
- databases;
- external systems.
The ambition is becoming more and more visible:
making Claude a permanent layer of work rather than a simple question-and-answer interface.

The benchmarks Anthropic highlights
Anthropic cites in particular the Online-Mind2Web benchmark, oriented towards web agents and navigation.
Announced result:
84%
Anthropic states that this score beats Opus 4.7 as well as GPT-5.5 on that specific benchmark.
As always, proprietary benchmarks should be read with caution.
But they clearly show where Anthropic is concentrating its efforts:
agents and autonomy.
The Opus 4.8 paradox
This is probably the most interesting point of the release.
On one side, Anthropic explains that it worked on the model’s honesty because AI systems:
- sometimes conclude too quickly;
- sometimes claim to have finished a task;
- can produce insufficiently verified information.
On the other:
Anthropic is simultaneously increasing:
- autonomy;
- task duration;
- the number of sub-agents;
- the capacity to act.
In other words:
the more the AI acts alone, the higher the potential cost of a mistake.
And the fact that “honesty” is now becoming a major marketing argument is probably not an accident.
That argument, incidentally, is what Anthropic would keep building on two weeks later when it opened a guarded version of its Mythos family to the public — a model that is more powerful and more supervised at the same time, and which hands requests back to Opus 4.8 when it judges them too sensitive.
What to remember
Claude Opus 4.8 is not the most spectacular update of the year.
But it may be one of the most revealing.
Because it shows where the industry is going.
The subject is no longer only:
- answer quality;
- benchmarks;
- scores.
The subject becomes:
how long an AI can work alone before a human needs to step in.
And Anthropic now seems to consider that question important enough to make the model’s honesty a central product argument.
The more autonomy AI gains.
The more its ability to recognise its own limits becomes strategic.
The figure to remember
1 million tokens of context, 128,000 tokens of output, hundreds of parallel sub-agents, and an AI its own maker now presents as more honest.
Anthropic’s message is clear:
the next battle is no longer only about how intelligent models are, but about their ability to work autonomously without making expensive mistakes.