On 22 September 2026, Anthropic released Claude Opus 5.5, the first model in a new Claude 5.5 family. A few hours later, the same day, OpenAI answered with GPT-6 Sol and GPT-6 Luna, halving its prices.
Both announcements tell the same story: almost as strong as the flagship, for far less money. Both target the same use, the one that burns the most tokens today — writing code. And both are sold with numbers attached.
The problem is that those numbers never meet. Anthropic compares Opus 5.5 to GPT-5.6 Sol, a model OpenAI replaced within the day. OpenAI compares GPT-6 Sol to Claude Opus 5 and Claude Fable 5, never to Opus 5.5. Neither one measured itself against its actual rival.
And the usual referee is missing: as of 23 September 2026, no independent coding leaderboard contains these two models. We checked the three that count, plus our own.
Here is what the only existing figures say, once recalculated.
The verdict, in short
Winner: Claude Opus 5.5. It wins the only two benchmarks both vendors publish in the same version, it wins the only independent measurement containing both models, and it wins on subscription price.
On measured intelligence: 58 against 48. On version 4.3.2 of the Artificial Analysis index, Opus 5.5 at maximum effort is first in the general leaderboard. GPT-6 Sol at its own maximum comes fifteenth.
On list price: GPT-6 Sol, half as much. $2 and $10 per million tokens against $4 and $20.
The figure that decides it. Opus 5.5 at its default setting scores 51 for $1.34 per task. GPT-6 Sol at its maximum scores 48 for $1.06. Three points more for 26% more money: the “expensive” model turned down beats the “cheap” model turned all the way up, and the gap on the bill is nothing like double.
The detail nobody mentions. Above 272,000 input tokens, OpenAI doubles its input rate. Anthropic applies no surcharge anywhere across its million-token window. On a whole code repository, input then costs exactly the same on both.
Where GPT-6 Sol genuinely wins: speed and volume. It starts replying in 2.0 seconds where Opus 5.5 takes 22.5, and running the entire index costs it $1,550 against $8,708.
The trap they share. Opus 5.5 at maximum effort is the worst buy in either catalogue: 4.5 times the price for seven index points.
Where the figures in this comparison come from
Three families of sources, never mixed.
The vendors. Prices, context windows, billing thresholds, effort levels, knowledge dates and access conditions come from the announcements of 22 September and the official documentation, recorded on 23 September 2026. Every vendor score is flagged as such, with the benchmark version and the setting used.
One independent measurement. The Artificial Analysis index, version 4.3.2. It is the only public record containing both models. This index is versioned: values published under an earlier version, including in our own previous comparisons, do not compare to these.
Our own recalculations. All costs were redone from the official pricing tables, with assumptions written out in black and white so you can replay them.
One warning, finally, that applies to everything below: token counts are not comparable from one vendor to the other. Anthropic itself writes in its pricing documentation that the tokeniser in its 4.7 and later models « produces roughly 30% more tokens for the same text » than the previous one. A price per token therefore does not, on its own, tell you what a given text will cost.
The two contenders
| Claude Opus 5.5 | GPT-6 Sol | |
|---|---|---|
| Vendor | Anthropic | OpenAI |
| Release date | 22 September 2026 | 22 September 2026 |
| Model ID | claude-opus-5-5 | gpt-6-sol |
| Input / output (per million tokens) | $4 / $20 | $2 / $10 |
| Cache read | $0.20 | $0.20 |
| Cache write | $5 (5 min) / $8 (1 h) | $2.50 |
| Context window | 1,000,000 tokens | 1,050,000 tokens |
| Long-context surcharge | none | ×2 input, ×1.5 output above 272,000 tokens |
| Maximum output | 128,000 tokens | 128,000 tokens |
| Effort levels | low, medium (default), high, xhigh, max | none, low, medium (default), high, xhigh, max |
| Knowledge date | June 2026 | 20 April 2026 |
| Batch pricing | $2 / $10 | not published on the model page |
Two remarks on that table.
First: cache reads cost exactly the same on both, $0.20 per million. That is not a coincidence of pricing. At OpenAI it is one tenth of the input price, the usual rule. At Anthropic it is a special multiplier of 0.05 — twice as generous as the house rule, and reserved for Opus 5.5. Anthropic explains why in its announcement: cache reads « represent the majority of costs for agentic and coding work ».
Second: Opus 5.5’s batch pricing, $2 and $10, is exactly GPT-6 Sol’s standard pricing. For any work that does not need an immediate answer — repository analysis, migration, background code review — the model reputed to be expensive costs the same as the model reputed to be cheap.
Nobody has measured them on code yet, and that is no accident
This is the point the comparisons published since yesterday pass over in silence. We looked for both models in the four leaderboards that carry weight. They are not there.
SWE-bench Verified, the historical coding reference, is closed to commercial labs. The official submissions repository carries a note dated 18 November 2025: from then on, only submissions meeting two cumulative conditions are accepted — a link to an open research publication, and at least one author affiliated with an academic institution or an established research lab. The reason is written plainly: keeping the benchmark as « a venue for advancing scientific understanding of code generation rather than a product validation platform ».
The effect is visible in the data. The repository contains 182 submissions. Only fourteen are from 2026: thirteen in February, one on 1 September. And none covers an Opus 5.x, GPT-6, Fable or Mythos model. The leaderboard still cited by most “best AI for coding” articles no longer measures the models they are writing about.
DeepSWE, maintained by Datacurve, is the agentic coding leaderboard we normally use. Updated on 22 September 2026, it stops at GPT-6 Astra (74% success, $4.43 per task), GPT-5.6 Sol (73%, $6.46), Claude Opus 5 (74%, $11.84) and Claude Fable 5 (70%, $13.41). Neither Opus 5.5 nor GPT-6 Sol.
compar:IA, the French state’s public ranking, fed by 264,029 blind votes as of 23 September 2026, contains no Opus or Fable model, and no GPT-6.
Our own ranking, finally. Regenerated on 22 September at 20:40 Paris time, it does contain both models at each of their effort settings — but with a still-partial evaluation, backed by a single source, and they appear in none of our code categories. So we will not quote a rank: that would be giving a precision the data does not have. Our ranking and comparison tool will update when the measurements arrive.
One conclusion to keep in mind for everything that follows: the only coding figures available on these two models are the ones their vendors chose to publish. Plus one independent general-purpose measurement, the Artificial Analysis index, which contains a coding benchmark but is not a coding leaderboard.
Neither compared itself to its actual rival
Look at the columns each vendor picked.
Anthropic’s table compares Opus 5.5 to Claude Fable 5.1, Claude Opus 5, GPT-6 Astra and GPT-5.6 Sol. That last one was replaced by GPT-6 Sol the same day, at half the price. The table’s most flattering comparison on the OpenAI side — CursorBench, where Opus 5.5 leads GPT-5.6 Sol by 11 points « for roughly a third of the cost per task » — therefore concerns a model no longer sold at that price.
OpenAI’s table compares GPT-6 Sol to GPT-6 Astra, GPT-5.6 Sol, GPT-5.6 Luna, Claude Opus 5 and Claude Fable 5 or 5.1. Never to Opus 5.5. OpenAI says so in a footnote: « competitor model evaluations come from publicly available reports », and « Claude Fable 5 scores were reported where Claude Fable 5.1 scores were not available ».
This is not bad faith, it is a calendar problem: you cannot measure a model that ships while you are putting your own press release online. But the result is the same for the reader. Both announcements of 22 September measure themselves against the past.
The only two benchmarks where the comparison holds
Across nine benchmarks published by Anthropic and six by OpenAI, only two carry the same name in the same version. Here they are.
| Benchmark | What it measures | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|
| FrontierCode 1.1 (main set) | would a patch be merged? | 54.4% | 49.3% |
| AutomationBench 1.0.6 | business workflows, 47 tools | 40.0% | 33.2% |
| Terminal-Bench 4.0 | command line | 66.4% | not published by OpenAI |
| DeepSWE v1.1 | long tasks, real repositories | not published by Anthropic | 68.8% |
| CursorBench 4.0 | multi-file tasks | 57.8% | not published by OpenAI |
Opus 5.5 wins both possible head-to-heads, by 5.1 and 6.8 points. But you have to read the footnotes, and they pull in both directions.
Anthropic writes that AutomationBench was run by Zapier with no fallback model, so every intervention by its safeguards counted as a failure — which, the vendor says, « produced a lower score than Opus 5.5 would achieve in practice ». The 40.0% would therefore be understated.
OpenAI levels exactly the opposite reproach at its competitor, on the same benchmark: the Claude Fable 5.1 data point « understates its true cost, as it omits the cost of fallbacks to Opus 5, which occurred on roughly 40% of tasks ». In plain terms: when Fable 5.1 could not cope, another model took over, and that second model’s bill was not counted.
Two vendors, two footnotes, one subject: these models do not work alone, and how the crutches are accounted for changes both the scores and the prices. We come back to this below.
One last point on that table: DeepSWE. GPT-6 Sol scores 68.8% there at maximum effort, which OpenAI presents as being « within 1.1 points of Claude Fable 5’s best score » — 69.9% — « at roughly 80% lower cost per task ». The figure is credible, but it compares Sol to Fable 5, not to Opus 5.5, and the official DeepSWE leaderboard has not verified it: it does not contain GPT-6 Sol.
The independent measurement: 58 against 48
One single set of measurements contains both models today: the Artificial Analysis intelligence index, version 4.3.2, ten benchmarks. Both models are measured there at maximum effort.
| Benchmark | What it measures | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|
| Terminal-Bench 4.0 | command line | 60% | 44% |
| SciCode | scientific code | 67% | 58% |
| AutomationBench-AA | business workflows | 70% | 62% |
| GDPval-AA v2.1 | real professional work | 1846 | 1487 |
| AA-Briefcase v1.1 | working through case files | 1822 | 1483 |
| AA-Omniscience | made-up answers | 46 | 27 |
| Humanity's Last Exam | expert questions | 61% | 48% |
| AA-LCR v1.1 | long context | 85% | 84% |
| GDP.pdf | document reading | 26% | 25% |
| CritPt | research physics | 32% | 31% |
Ten benchmarks, ten times Opus 5.5 ahead or level. Composite score: 58 against 48.
But the distribution matters as much as the total. On the last three rows — long context, document reading, physics — the gap is one point. In other words, where the job is mostly to read a lot and produce an answer, the two models are level, and paying double makes no sense.
The gap opens up where the model has to act in several steps: 16 points on the terminal, 8 on automation, 359 Elo points on real professional work. That is exactly the profile of a coding agent that opens files, runs tests, reads the errors and starts again.
One row deserves a pause: AA-Omniscience, 46 against 27. This benchmark rewards correct answers, penalises made-up ones and does not punish abstention. For code, a made-up answer is expensive: a function that does not exist, an invented parameter, a phantom dependency. On that specific point, the gap is close to double.
Finally, the general leaderboard recorded on 23 September 2026 places Opus 5.5 at maximum effort first in the world, ahead of Claude Fable 5.1 and GPT-6 Astra. GPT-6 Sol at its maximum is fifteenth.
The figure that decides it: the default setting beats the other’s maximum
This is the most useful calculation in this comparison, and it inverts the intuition.
| Setting | Index | Cost/task | Tokens |
|---|---|---|---|
| Opus 5.5, maximum effort | 58 | $5.98 | 119,000 |
| Opus 5.5, default setting | 51 | $1.34 | 25,700 |
| GPT-6 Sol, maximum effort | 48 | $1.06 | 31,200 |
Read the middle row. Opus 5.5 left at its factory setting does better than GPT-6 Sol pushed all the way, for 26% more per task. Not double: 26%.
And it writes fewer tokens than its competitor — 25,700 against 31,200 — while scoring three points higher. The model reputed to be verbose is here the more economical of the two by volume.
Two practical consequences.
First: the question “which one is cheaper” has no answer until you fix the effort setting. Comparing a price per token without saying what effort each model is running at is comparing two cars without saying how fast they are going.
Second: if budget is your first criterion, the right choice is not necessarily GPT-6 Sol — it may be Opus 5.5 turned down. Anthropic publishes several customer accounts pointing the same way. Deloitte reports that « even at its lowest effort setting », Opus 5.5 caught 72% of known bugs in its code reviews, against 56% for Opus 5 at high effort. Rogo says the same about finance: Opus 5.5 at minimum effort beating Opus 5 at high effort, with « roughly 60% fewer output tokens ». These are customer figures relayed by the vendor, not independent measurements — but they all point in the same direction.
The maximum-effort trap
Now look at the first row of the same table. $5.98 per task for an index of 58, against $1.34 for 51. Seven index points cost 4.5 times the price, and 4.6 times more tokens written.
Worse: on at least one coding benchmark, turning the effort up buys nothing. On FrontierCode, Anthropic records 54.6% at the default setting and 54.4% at maximum effort. Within the margin of error, certainly — but it means the premium buys nothing at all.
The vendor itself warns about this, which is rare enough to quote. In its announcement, Anthropic writes that « at these capability levels, benchmark gaps have become a less reliable guide to real-world differences », and that « in our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest ».
Simple rule: set it low, and turn it up only when you actually see a failure. Not the other way round.
Half price, up to 272,000 tokens
Here is the detail that appears in neither announcement and that directly concerns anyone working on real code repositories.
GPT-6 Sol’s developer page contains a discreet line: « requests above 272,000 input tokens are billed at double the input and cache rates, and one and a half times the output rate ».
Anthropic’s pricing documentation says exactly the opposite for its 4.6 and later models: the full million-token window is billed at the standard rate, and the example is explicit — « a 900,000 token request is billed at the same per-token rate as a 9,000 token request ».
Which gives this:
| Request | GPT-6 Sol | Opus 5.5 | Gap |
|---|---|---|---|
| Up to 272,000 tokens | $2 / $10 | $4 / $20 | double |
| Above 272,000 tokens | $4 / $15 | $4 / $20 | input identical |
Take a concrete, checkable case. You send a repository of 400,000 tokens — a medium-sized project, or a large module with its history — and the model hands you back 20,000 tokens of code and explanation.
- GPT-6 Sol: 400,000 × $4 / 1,000,000 = $1.60 input, plus 20,000 × $15 / 1,000,000 = $0.30 output. Total: $1.90.
- Claude Opus 5.5: 400,000 × $4 / 1,000,000 = $1.60 input, plus 20,000 × $20 / 1,000,000 = $0.40 output. Total: $2.00.
Five per cent apart. Not one hundred per cent. And that is precisely the codebase-work scenario — migration, audit, cross-cutting review — that both vendors put forward in their announcements.
The threshold itself is not trivial either: 272,000 tokens is about a quarter of the window OpenAI advertises. GPT-6 Sol’s million tokens exist, but the top three quarters are billed at Opus 5.5’s rate.
Cache costs exactly the same on both
A second correction to “half price”.
In an agentic coding session, most of what is sent to the model is re-read, not discovered: the same files, the same instructions, the same history, on every turn. That is what caching is for, and Anthropic writes explicitly that cache reads « represent the majority of costs for agentic and coding work ».
Both bill cache reads at $0.20 per million tokens. At OpenAI that is one tenth of the input rate. At Anthropic it is an exceptional multiplier of 0.05, against 0.1 on all its other models outside Fable and Mythos.
On the heaviest line of a coding session’s bill, the price gap between the two models is zero.
That leaves cache writes, where GPT-6 Sol is cheaper: $2.50 per million against $5 at Anthropic for a five-minute cache and $8 for a one-hour cache. But one write serves several reads: Anthropic states that a five-minute cache pays for itself on the first re-read, and a one-hour cache on the second.
OpenAI, for its part, worked on the mechanics rather than the price: better cache hit rates by default, a monitoring dashboard, a diagnostic tool, explicit breakpoints, and above all the ability to change the effort level or the tool list without breaking the cache. GitHub reports that these improvements cut « by more than 50% » the share of tokens needing fresh processing, across billions of requests.
The bill for a day of coding, recalculated
Let us set out simple, written assumptions, so you can redo them.
Assumptions. One working day, 40 agent tasks — roughly one every twelve minutes over eight hours. Cost per task as measured by Artificial Analysis on its index, which mixes code, agent work, reading and reasoning: it is an approximation of real use, not a bill for your project.
| Setting | 1 day (40 tasks) | 20 days | Index |
|---|---|---|---|
| GPT-6 Sol, maximum effort | $42.40 | $848 | 48 |
| Opus 5.5, default setting | $53.60 | $1,072 | 51 |
| Opus 5.5, maximum effort | $239.20 | $4,784 | 58 |
What this table says. Moving from GPT-6 Sol at maximum to Opus 5.5 at default costs $224 more a month and buys three index points. Moving from Opus 5.5 at default to Opus 5.5 at maximum costs $3,712 more a month for seven points.
The first trade-off is arguable. The second, for a single developer, is not.
For scale: OpenAI writes in its announcement that its own researchers’ daily token consumption, valued at its API rates, exceeds $600 for the median researcher and $7,000 at the 90th percentile. Our 40 tasks a day is, if anything, conservative for heavy use.
And if you do not need an immediate answer — overnight analysis, migration, deep review — Anthropic’s batch pricing brings Opus 5.5 down to $2 and $10, which is GPT-6 Sol’s standard rate. In that mode, the price debate disappears entirely.
Speed: 2 seconds against 22
This is GPT-6 Sol’s clearest win, and it shows up in no score.
| Measurement | GPT-6 Sol | Opus 5.5 |
|---|---|---|
| Tokens per second (medium effort) | 114.3 | 75.8 |
| Time to first token (medium effort, 10,000 token input) | 2.0 s | 22.5 s |
| Time to first token (maximum effort) | — | 149.29 s |
| Cost to run the whole index | $1,550 | $8,708 |
Eleven times faster to start. For a developer working conversationally — ask, read, correct, ask again — twenty seconds of waiting before the first character, on every turn, changes the experience completely. And at maximum effort, Opus 5.5 thinks for nearly two and a half minutes before writing anything at all.
Anthropic offers a partial answer: a fast mode, in research preview, up to 2.5 times quicker, billed at $8 and $40 per million tokens — double the normal rate, and four times GPT-6 Sol’s. It is available in Claude Code and on Anthropic’s own direct API, but not through Amazon Bedrock or Google Cloud, and not with batch processing.
The last row of the table shows the other face of the same phenomenon: running the entire index costs 5.6 times more on Opus 5.5. For high-volume, low-stakes use — classification, log summarisation, repetitive test generation — that gap is decisive, and GPT-6 Sol is the right choice. For that kind of task, GPT-6 Luna at $0.10 and $0.50 is better suited still.
The crutches each vendor throws at the other
Neither of these models works entirely alone, and that is something the score tables hide poorly.
Anthropic writes it in the notes to its own table: Opus 5.5 « was evaluated with its production protections enabled. When they intervened, cybersecurity tasks were carried out by Claude Opus 4.8, and biology and frontier model development tasks by Claude Opus 5 ». And it adds that this « likely reduces » the performance shown for Opus 5.5.
In other words: on some rows of that table, it was not Opus 5.5 answering. Artificial Analysis draws the consequence in its own leaderboard, where Claude entries carry the explicit label « with default fallback ».
OpenAI turns the same argument back on Anthropic, this time on cost. In its AutomationBench note: the Claude Fable 5.1 data point « understates its true cost, as it omits the cost of fallbacks to Opus 5, which occurred on roughly 40% of tasks ».
Forty per cent. If a score announced for one model was obtained four times out of ten by another, more expensive model, then neither the score nor the price shown describes what you are buying.
What to take from this: when a vendor publishes a score, ask yourself which model actually did the work, and who pays the difference. It is the same reflex we apply across our AI comparisons and our tested promises.
What each one knows about your language
For code, the knowledge date is not a detail: it decides whether the model knows the version of the library you are using, or offers you the interface from two versions ago.
- Claude Opus 5.5: reliable knowledge date and training date set to June 2026.
- GPT-6 Sol: 20 April 2026.
Two months apart, in Anthropic’s favour. Not a chasm, but in an ecosystem where a major framework ships several versions a quarter, it is enough to create silent errors — code that compiles and uses a deprecated method.
Both models have web search to fill the gap case by case. At Anthropic it is billed at $10 per 1,000 searches on the API, on top of the tokens; page retrieval costs only the tokens. Neither vendor publishes the list of libraries actually covered.
A word on the windows while we are here. GPT-6 Sol advertises 1,050,000 tokens, of which 922,000 maximum as input. Opus 5.5 advertises 1,000,000. On paper OpenAI is ahead; on the bill, we have seen what happens above 272,000. And for output, both cap at 128,000 tokens — Anthropic goes up to 300,000 on its batch API, with a dedicated header.
If you work in cybersecurity, Opus 5.5 is not the one answering
This point deserves its own heading, because it affects an entire population of developers.
Anthropic writes that Opus 5.5 is « comparable to Claude Mythos 5.1 in biology and cybersecurity », and that it is therefore deployed with safeguards equivalent to those of Fable 5.1. The concrete consequence: cybersecurity tasks are routed to Claude Opus 4.8, a previous-generation model.
A verification programme is announced for practitioners in the field, with « three increasingly permissive access tiers », but with no availability date. In the meantime, if your job is writing detection rules, analysing a binary or auditing an exploit chain, you will not get answers from the model you are paying for.
GPT-6 Sol documents no equivalent mechanism in its announcement. OpenAI instead highlights progress on what it calls coding deception: the model less often claims to have done work it has not done. Five evaluation categories are cited — coding deception, broken search, reviewer bypass, warning bypass, unauthorised interaction — with no public figures in the announcement, which refers to the system card.
On resistance to hidden malicious instructions, Anthropic states that Opus 5.5 « matches or exceeds Opus 5 in all tested settings », and that it ties with Fable 5.1 for the lowest injection success rate of all models tested by the security firm Gray Swan.
On subscriptions: 15 euros against 23
Most developers do not call these models through the API: they go through a subscription. The official tables, recorded on 23 September 2026, favour Anthropic. Prices below are the European rates.
| Plan | Price | Gives access to |
|---|---|---|
| Claude Free | 0 € | Sonnet and Haiku. Neither Opus nor Claude Code |
| Claude Pro | 15 €/month annually (180 € up front), 18 € monthly | Opus 5.5, Sonnet, Haiku, Claude Code included |
| Claude Max | from 90 €/month | 5× or 20× Pro usage, priority access |
| ChatGPT Free | 0 € | GPT-5.6 Luna unlimited in text, Codex « limited ». Not GPT-6 Sol |
| ChatGPT Go | 8 €/month | More messages. Not GPT-6 Sol, and ads possible |
| ChatGPT Plus | 23 €/month | GPT-6 Astra, Sol and Luna, extended Codex usage |
| ChatGPT Pro | from 103 €/month | 5× more usage, maximum Codex tasks |
The entry ticket is 15 to 18 euros at Anthropic, against 23 at OpenAI, to reach the model in question here. Claude Code is included in all of Anthropic’s paid plans and shares the same counter as conversations: your terminal work and your chats draw on the same allowance.
Two details that cost money if you miss them.
At Anthropic, access to Fable 5.1 — the house’s strongest model, at $10 and $50 per million — is not included in Pro: it consumes usage credits. On Max plans it is capped at « 50% of weekly limits ». Anthropic also announces, alongside Opus 5.5, higher five-hour limits for Pro, Max, Team and Enterprise, and a quota reset « that you can keep and use whenever you want ».
At OpenAI, GPT-6 Sol is available « in ChatGPT Work and Codex » for Plus, Pro, Business, Enterprise and Edu — but the announcement specifies that it is not yet available in Chat. The pricing page also mentions a temporary offer doubling Codex usage, announced until 31 May 2026.
Your data, Europe and watermarking
Three concrete points, all taken from the official documentation.
Zero retention. Opus 5.5 is available with a zero data retention option, like previous Opus models.
Watermarking and the AI Act. Anthropic states that Opus 5.5, like Fable 5.1, ships with its « watermarking measures to comply with the European AI Act ». Services in the European Union are provided by Anthropic Ireland Limited.
Data residency. On Claude 4.6 and later models, requesting inference exclusively in the United States applies a 1.1 multiplier to all price categories, input, output, cache writes and cache reads included. Global routing, which is the default behaviour, stays at the standard rate. Conversely, regional and multi-regional endpoints at Amazon and Google cost 10% more than global endpoints.
One last change can break existing integrations without warning: Anthropic flags a modification to how retained reasoning is handled, applying to Fable 5.1 and Opus 5.5 « for accounts created from 31 August 2026 ». If you built an agent on a recent account, check that before switching.
And in French?
No public measurement settles it, and that is worth saying rather than improvising an answer.
Neither Anthropic’s nor OpenAI’s announcement publishes a French-language evaluation for these two models. On compar:IA, the French state’s public ranking fed by blind votes from French-speaking users, the record of 23 September 2026 covers 68 models and 264,029 votes: neither Opus 5.5 nor GPT-6 Sol appears, and no Opus or Fable model is ranked. The closest reference points are GPT-5.6 Terra at 1,120 points and GPT-5.6 Luna at 1,104.
OpenAI does mention work on communication style — « more clarity, less jargon, fewer odd turns of phrase, slightly shorter answers » — and Anthropic quotes a tester saying of Opus 5.5 « it writes like me ». Both statements concern English and neither is quantified.
For code, the French question arises mainly for comments, commit messages and natural-language exchanges with the agent. Nobody measures it.
Which one to choose
Take Claude Opus 5.5 if you write code seriously, on an existing codebase, with an agent that opens files, runs tests and starts over. It wins both available head-to-heads, it wins the only independent measurement, it is first in the general leaderboard, and its default setting is enough to beat GPT-6 Sol pushed all the way. Set it to medium effort and turn it up only when you see a failure. Its entry subscription is also the cheaper one, at 15 euros a month annually.
Take GPT-6 Sol if latency matters more to you than the last point of quality — two seconds against twenty-two before the first character — if your volume is high and your tasks repetitive, or if you need a « none » effort level that Claude no longer offers. It is also the default choice if you are already inside Codex and switching tools would cost more than the difference between the models.
Take neither at maximum effort without checking that you need it. That is true of both, but it is spectacular at Anthropic: 4.5 times the price for seven index points, and zero measurable gain on FrontierCode.
Do not pay for long context you are not using. And if you really are using it, know that above 272,000 tokens the price gap between the two drops to 5%.
Wait a week or two before locking in your choice, if you can. The independent coding leaderboards will add these models, Anthropic announces Sonnet 5.5 and Haiku 5.5 « in the coming weeks », and our own measurements will be updated at the same pace.
What to remember
On 22 September 2026, hours apart, Anthropic released Claude Opus 5.5 and OpenAI answered with GPT-6 Sol at half the price. Both target code.
No independent coding leaderboard has measured them yet. SWE-bench Verified stopped accepting commercial labs on 18 November 2025; DeepSWE, compar:IA and our own ranking do not contain them in their code categories as of 23 September.
Neither vendor compared itself to the other. Anthropic measures itself against GPT-5.6 Sol, replaced the same day. OpenAI measures itself against Claude Opus 5 and Fable 5.
On the only two common benchmarks, Opus 5.5 wins: 54.4% against 49.3% on FrontierCode 1.1, 40.0% against 33.2% on AutomationBench. On the only independent measurement it wins too: 58 against 48, and first place in the general leaderboard.
The figure that decides it: Opus 5.5 at its default setting (51) beats GPT-6 Sol at its maximum (48), for $1.34 per task against $1.06. Twenty-six per cent more, not double.
Above 272,000 tokens, the price gap collapses to 5%, because OpenAI doubles its input rate and Anthropic applies no surcharge. And cache reads, which carry most of a coding agent’s bill, cost exactly $0.20 on both.
GPT-6 Sol keeps two real advantages: it starts eleven times faster, and it costs 5.6 times less to run at volume.
Winner for writing code: Claude Opus 5.5, at its default setting. With a caveat the vendor states itself: at this level, « benchmark gaps have become a less reliable guide to real-world differences ».