Tool

AI model comparator

Pick three models, get the verdict use by use. No test to run, no prompt to write — the comparison is already calculated.

  1. 1Tick three modelsin the table below, or in any ranking on the site.
  2. 2Run the comparisonin one click. Nothing to install, nothing to create.
  3. 3Read the verdictuse by use: who wins, by how much, and how solid the measurement is.

An example, calculated today

Claude Opus 4 7 HighAnthropic vsGemini 3 ProGoogle vsGrok 4.20 Beta 0309 ReasoningxAI

Claude Opus 4 7 High leads 5 uses · Gemini 3 Pro comes first on none · Grok 4.20 Beta 0309 Reasoning comes first on none.

5 usages départagés sur 5 comparables · 8 usages hors périmètre

Use-by-use comparison
Use Claude Opus 4 7 HighGemini 3 ProGrok 4.20 Beta 0309 Reasoning Winner
Généraliste 96,0 #3 92,1 #10 90,7 #18 Claude Opus 4 7 Highcaveat
Conversation 96,0 #16 93,0 #29 90,7 #50 Claude Opus 4 7 High
Code 61,4 #66 46,1 #207 37,4 #278 Claude Opus 4 7 High
Rapport coût / performance 96,0 #20 93,0 #31 90,7 #45 Claude Opus 4 7 High
Agents cloud 96,0 #2 92,1 #9 90,7 #13 Claude Opus 4 7 Highcaveat

The comparison is recalculated on every ranking update. Pick your own models below.

Compare the models of your choice

Comparateur Le Recul

Who wins on what

Cochez au moins deux modèles dans le tableau du classement, puis revenez ici.

This list shows only the top twenty of the general ranking. The remaining 1,800 models — including code, image, video, audio and open models — are compared the same way from the full ranking, where the comparator works identically in each of the 13 categories.

Top twenty models in the general ranking
# Model Vendor Score
01 Claude Opus 4 6 High Anthropic 96,5 / 100
02 Claude Opus 5.5 High Anthropic 96,3 / 100
03 Claude Opus 4 7 High Anthropic 96,0 / 100
04 Claude Opus 4.6 Anthropic 95,1 / 100
05 Claude Fable 5 High Anthropic 94,7 / 100
06 Claude Opus 4.7 Anthropic 94,6 / 100
07 Claude Opus 5 High Anthropic 94,1 / 100
08 Claude Opus 4 8 High Anthropic 92,3 / 100
09 Claude Fable 5 1 Max Anthropic 92,2 / 100
10 Gemini 3 Pro Google 92,1 / 100
11 GPT 5.2 Chat Latest 20260210 OpenAI 91,4 / 100
12 Claude Opus 4.8 Anthropic 91,3 / 100
13 GPT-5.4 high OpenAI 91,3 / 100
14 Grok 4.20 beta1 xAI 91,3 / 100
15 Qwen3.7 Max Preview Alibaba 91,3 / 100
16 Claude Opus 4 5 20251101 High 32k Anthropic 91,1 / 100
17 Claude Sonnet 4.6 Anthropic 90,7 / 100
18 Grok 4.20 Beta 0309 Reasoning xAI 90,7 / 100
19 Grok 4.20 Multi Agent Beta 0309 xAI 90,6 / 100
20 Gemini 3 flash Google 90,2 / 100

Compare across all 1,800 models →

Our written comparisons

The tool calculates. Our written comparisons explain: context, limits, and what the figures do not say.

Best AI for coding: Claude Opus 5.5 against GPT-6 Sol, released the same day hours apart

Comparisons

Best AI for coding: Claude Opus 5.5 against GPT-6 Sol, released the same day hours apart

On 22 September 2026 Anthropic released Claude Opus 5.5, and OpenAI answered the same day with GPT-6 Sol at half the price. Both are sold for writing code. No independent coding leaderboard has measured either of them: SWE-bench Verified stopped accepting commercial labs in November 2025, and the DeepSWE leaderboard does not contain them. We went through both announcements line by line, cross-checked the only independent measurement available, recalculated every bill, and read the pricing tables down to the thresholds. Three things come out of it that appear in neither announcement: the more expensive model at its default setting beats the cheaper one at maximum, the price gap collapses exactly where code needs it, and neither vendor compared itself to the other.

DeepSeek V4.1 Flash vs GPT-5.6 Luna: the cheaper of the two depends on what time you call it

Comparisons

DeepSeek V4.1 Flash vs GPT-5.6 Luna: the cheaper of the two depends on what time you call it

DeepSeek released V4.1 Flash on 10 September 2026, and the price everyone quotes is striking: $0.15 per million tokens. Facing it, GPT-5.6 Luna, the model behind free ChatGPT and the entry point of OpenAI's GPT-5.6 family. On the independent index, DeepSeek edges ahead, 40 to 38. But that price is the off-peak rate, and it is billed in UTC — which changes the answer completely depending on where your team sits. DeepSeek also writes twice as many tokens for the same task, and its record coding score still has not been verified by the official leaderboard. We recalculated every bill, hour by hour.

Updated
Gemini Omni 1.1 Flash VS Wan 3.0: the better of the two is half switched off, and which half depends on where you are

Comparisons

Gemini Omni 1.1 Flash VS Wan 3.0: the better of the two is half switched off, and which half depends on where you are

Google shipped Gemini Omni 1.1 Flash on 27 August 2026; Alibaba had opened Wan 3.0 a few weeks earlier. Thirty-five hundredths of a point separate them in our ranking. We read both announcements word for word, recalculated the cost of a delivered minute with the discarded takes included, gathered the reports of people running them in production, and read the documentation down to the restrictions the vendors write in small print. What comes out is in neither press release: Google's forty seconds are ten-second slices stacked end to end, its 4K is an upscale, and the feature that makes the model interesting — editing video you shot yourself — is switched off in the United Kingdom and across the European Economic Area, at API level. Meanwhile the Chinese model generates thirty seconds in one take, accepts a PDF as input, and publishes its own list of weaknesses. And one criterion flips for an English-speaking reader.

See all comparisons →

Frequently asked questions

How many models can be compared at once?

Three. Beyond that the table becomes unreadable and the comparison loses its point: you do not choose between eight models, you choose between two or three finalists.

Where do the scores come from?

From cross-checked public sources, aggregated daily. Every score carries its evaluation status and the number of independent sources backing it.

What does “too close to call” mean?

That the gap between two models is smaller than the measurement uncertainty. Naming a winner in that case would be an arbitrary choice presented as a result: the comparator would rather say so.

Why do some cells say “not measured”?

Because the model does not appear in any public source for that use. Nothing is estimated or filled in: an empty cell is information, not an oversight.

Do I need an account?

No. The comparator is free, with no sign-up and no account. No personal data is requested.