Ask ten people “which AI should I use” and you will get ten answers, all sincere and all partial. The one who codes will swear by one model, the one who writes by another, the one who generates images by a third. They are all right, and that is precisely the problem: there is no best AI, there is a best AI for what you want to do with it.
That is why we built an AI model comparator. You tick up to three models, you click, you get the verdict use by use: who wins, by how much, and with what measurement strength. It is free, with no sign-up, no account. And it does something tool comparisons almost never do: when the gap is smaller than the margin of error, it refuses to name a winner.
Here is how to use it, what each figure means, where the data comes from — and what the tool cannot do.
Which AI to use? That is not quite the right question
A single example is enough to break the idea of a single ranking. Here are three very well-known models, recorded in our ranking on 18 September 2026, with their score out of 100 and their rank in each category.
| Use | Claude Fable 5 | GPT-5.4 high | Gemini 3 Pro |
|---|---|---|---|
| General | 100.0 — 1st | 94.6 — 9th | 94.2 — 11th |
| Chat | 100.0 — 1st | 94.6 — 27th | 96.2 — 20th |
| Code | 75.0 — 18th | not measured | 48.3 — 179th |
| Cost against performance | 100.0 — 8th | 94.6 — 31st | 96.2 — 26th |
| Cloud agents | 100.0 — 1st | not measured | 94.2 — 8th |
Look at the highlighted row. The same model, Gemini 3 Pro, is 20th in chat and 179th in code. A reader who had picked their subscription on the general ranking alone — where the gap between 9th and 11th fits within four tenths of a point — would have got it half wrong if what they mainly wanted was to code.
That is why “which is the best AI” is a question with no useful answer, and “which AI for which use” has one.
The comparator in one minute
Before the detailed guide, here is the tour. This one-minute video is the one that presents the tool on the site; the home page runs a short seventeen-second version, in a banner that will clear itself at the start of October. This article stays.
In short, the tool covers 1,892 models as recorded on 18 September 2026 and thirteen uses. It asks for no account, no email address, no payment. The scores are recalculated once a day, and every comparison shows the date of the data it uses.
The thirteen uses are not window dressing: they correspond to distinct rankings, fed by different sources.
| Family | Uses compared |
|---|---|
| General uses | General · Chat · Code · Cost against performance |
| Autonomous agents | Cloud agents · Coding agents · Local agents |
| Creation | Image · Video · Audio |
| Open and local models | Open source / open weights · Local — chat · Local — code |
How to use it: tick, compare, read
It all comes down to three gestures, none of which requires technical knowledge.
1. Tick up to three models. The checkbox is at the far left of each row, in the comparator table as in any table of the AI ranking. A bar appears at the bottom of the screen summarising your selection.
2. Click “Compare”. Nothing to install, nothing to create, no prompt to write: the comparison is already computed server side, it appears immediately.
3. Read the verdict, use by use. That is the part worth pausing on, and the subject of the next three sections.
Two details that save time:
- your selection follows you from page to page during your visit. You can tick a model in the Chat ranking, another in the Code ranking, a third in the Image ranking, then launch the comparison: all three are still there;
- the “Copy table” button, at the foot of the panel, copies the comparison as text. Handy for pasting into a message, a report or a team discussion.
On a phone, the tool works identically: the panel fills the screen, the figures stay readable, nothing is cut.
The verdict: who wins, by how much, and how solidly
The panel opens on a plain sentence — “X leads on 5 uses, Y comes first on none” — followed by the count of uses decided, then a chip per use with the winner’s name.
The most interesting line is the last one: “With caveats”. It appears when the model leading on a use is less well measured than the one it beats — fewer independent sources document it. The gap exists, but it rests on a more fragile base, and the tool says so rather than displaying a clean win.
It is an editorial choice with a cost: it makes the verdict less spectacular. It does, however, avoid turning a measurement uncertainty into a purchase recommendation.
Use by use: the table that also shows the holes
Further down the panel, the full detail: one row per use, each model’s score and rank, a gap bar, and the “Who wins” column.
Three notes recur, and each means something precise:
- “not measured”: no public source covers that model for that use. The cell stays empty, nothing is estimated or extrapolated;
- “Not applicable”, at the foot of the table: the use applies to none of the selected models — comparing three cloud models in the “Local agents” category would make no sense, so the row is dropped;
- “caveat”: the win stands, but on a narrower measurement base, as explained above.
A switch just above lets you show only what decides: tied and out-of-scope rows disappear, leaving only the uses where a model really takes the advantage. When you are hesitating between two subscriptions, that is often the only view that counts.
When the comparator refuses to decide
Here is the screen we are proudest of, and it is the one that shows the fewest results.
On that trio, the three general scores of 18 September 2026 sat within four tenths of a point: 94.6 against 94.2 against 94.4. Three different rankings measure them neck and neck. Naming a winner in those conditions would mean presenting a rounding as a result.
The comparator therefore shows “Too close to call”, and adds the sentence that sums up our position: a tie is a result, not a gap in the data. It says the choice is not decided by the score, but by what each model can do, by its price and by its availability — three criteria no benchmark measures for you.
The panel also shows, just above the detail, the “Consumer subscription” line: on 18 September 2026, 23 euros a month for the Anthropic and OpenAI offers quoted here, 22 euros a month for Google’s, and “Free” for Grok 4.20 beta1. When the scores are level, price becomes the deciding argument again.
Image, video, audio, local AI: the same tool everywhere
The comparator is not reserved for chat assistants. It works identically in each of the ranking’s thirteen categories, including for image generators, video models, audio, and open models that run on your own machine.
This screenshot shows one last function, useful when the figures are no longer enough: when we have written a full comparison of the selected models, the panel says so and gives the link. Here, 99.4 against 99.1 in video: a statistical tie, but our article explains that the better of the two is half disabled in Europe — a difference any score would have missed.
The same principle applies elsewhere. For a free AI, our comparison of Muse Glimmer against free ChatGPT sets out what you really lose without a subscription. For an AI that runs on your own computer, without sending your documents to the cloud, OpenClaw against Hermes compares two local agents on Windows. The comparator gives the figure; the article gives the context.
Three concrete situations, and where to start
The comparator gives figures; you still have to know which ones to look at. Here is how we would use it ourselves in three common cases.
You are starting out and only want to pay for one subscription. Start with the Chat category, not the general ranking: it is the one that measures what you will do every day — writing, summarising, explaining, rephrasing. Tick the two or three names you already know, run the comparison, and look first at the “Consumer subscription” line. If the verdict announces a tie, the question is no longer “which is best” but “which costs least and which do I prefer using”. A free model that suits you beats a paid model three tenths of a point better.
You code, or you have people code. The general ranking is the worst possible guide: our reading of 18 September 2026 gave a gap of 100 to 48.3 between two models separated by just two points in general use. Go straight to the Code ranking, then look at the Coding agents category, which measures something else: the ability to see a complete task through, not just to write a correct function. The two rankings do not give the same order, and that is normal — they do not ask the same question.
You are deciding for a small business. Three criteria weigh more than the raw score. Cost against performance, which relates results to the public subscription price — the only one you will pay if you equip five people. Cloud agents, if you want to hand over whole tasks rather than ask questions. And the Local and Open source categories, if your documents must not leave your machines: a model running on your own computer will never top the general ranking, but it answers a constraint the leaders cannot meet. In that case, compare local models with each other, never a local model with a cloud model — the tool automatically drops uses that apply to none of the chosen models anyway.
In all three cases the reflex is the same: pick the category matching your real use, tick two or three finalists, and only open the general ranking last — never first.
Where the figures come from
The score shown is not an in-house mark pulled out of a hat, and nor is it one ranking repeated. It is a composite score, computed from public benchmarks, normalised to 100 and weighted by category:
- LMArena, where users compare two answers blind and vote — the most widely used measure of human preference;
- SWE-bench, which asks a model to fix real bugs in real code repositories;
- Aider, in Polyglot and architect modes, for code editing in real conditions;
- LiveBench, whose questions are refreshed to limit contamination of training sets;
- MMLU-Pro, from the TIGER-Lab, for knowledge across fourteen disciplines;
- BBEH, published by Google DeepMind, for hard reasoning;
- Artificial Analysis, when the connection is live, for performance, price and speed.
Each model carries an evaluation status: Reliable when at least two independent sources confirm it, or when one reference source alone covers at least 60% of the category’s weighted criteria; Partial when a single source documents it — often a recent model, not yet cross-checked; Insufficient when the data does not allow a clean verdict. That status is not decorative: it is what triggers the “With caveats” note in the verdict.
Two points matter for daily use. First, the Cost against performance category compares performance with the price of the consumer subscription — not with API prices billed per million tokens, which do not reflect what an individual pays. Second, if a source fails to answer during a cycle, the ranking is not emptied: the last reliable measurement is kept within a freshness limit, and a safeguard blocks any publication that would abruptly degrade the data. The full methodology, including the per-category weightings, can be expanded at the foot of the AI ranking page.
What the comparator does not tell you
An honest tool is one that states its blind spots. Here are four, without evasion.
It does not test the models itself. It aggregates public measurements. A benchmark measures what it knows how to measure, not your job. The tests we run ourselves live in the Tested promises section, and they often tell a different story from the scores.
It does not measure any particular language. Most public benchmarks are in English. A model that is excellent in English can be less comfortable writing a summary note in French, German or Italian, and no column will say so.
It measures neither the interface, nor usage limits, nor privacy. The number of messages a day, the quality of the mobile app, what the vendor does with your data: in real life those criteria often weigh more than three tenths of a point.
The scores move. All the figures quoted in this article are those of 18 September 2026. The ranking is recalculated daily; a model newly covered by a second source can gain ten places overnight without anything having changed in the model itself. That is why every comparison shows its date.
What to take away
If you keep only four sentences:
- The best AI does not exist: the same model can be 20th on one use and 179th on another.
- Three models, thirteen uses, one verdict: the comparator answers in thirty seconds, free and with no sign-up.
- A tie is a result. When the gap falls below the measurement margin, the tool says so — and the choice then shifts to price, features and availability.
- Empty cells are information. A “not measured” is better than an invented figure.
The rest is yours: tick your two or three finalists, look at the use that matters to you, and drop the general ranking.
Our verdict. AI makes promises, Le Recul verifies: we will not tell you which AI is best, because nobody honestly can. We tell you which one wins on the use you care about, by how much, from which sources — and when the answer is “it is too close to call”.
Try it now: the AI model comparator is open, free and sign-up free. Three models, thirteen uses, one dated verdict.