AI ranking
Chat
User preference, reasoning, factuality, long context, simplicity, cost, speed.
Top visuel
- 01 Claude Opus 4 6 High 96.50
- 02 Claude Opus 5.5 High 96.33
- 03 Claude Opus 4 7 High 95.98
- 04 Claude Opus 4.6 95.10
- 05 Claude Fable 5 High 94.71
- 06 Claude Opus 4.7 94.58
- 07 Claude Opus 5 High 94.06
- 08 Claude Opus 4 8 High 92.31
- 09 Claude Fable 5 1 Max 92.23
- 10 Gemini 3 Pro 92.07
Compare models. Cochez jusqu'à 3 modèles dans le tableau ci-dessous, puis lancez la comparaison : qui gagne sur quel usage, de combien, et avec quelle solidité.
| # | Model | Vendor | Score | Monthly subscription | Assessment | |
|---|---|---|---|---|---|---|
| 01 | Anthropic | 96.50 / 100 | 23 €/mois | Partial | ||
| 02 | Anthropic | 96.33 / 100 | 23 €/mois | Partial | ||
| 03 | Anthropic | 95.98 / 100 | 23 €/mois | Partial | ||
| 04 | Anthropic | 95.10 / 100 | 23 €/mois | Partial | ||
| 05 | Anthropic | 94.71 / 100 | 23 €/mois | Reliable | ||
| 06 | Anthropic | 94.58 / 100 | 23 €/mois | Partial | ||
| 07 | Anthropic | 94.06 / 100 | 23 €/mois | Partial | ||
| 08 | Anthropic | 92.31 / 100 | 23 €/mois | Partial | ||
| 09 | Anthropic | 92.23 / 100 | 23 €/mois | Reliable | ||
| 10 | 92.07 / 100 | 22 €/mois | Reliable | |||
| 11 | OpenAI | 91.43 / 100 | 23 €/mois | Partial | ||
| 12 | Anthropic | 91.26 / 100 | 23 €/mois | Partial | ||
| 13 | OpenAI | 91.25 / 100 | 23 €/mois | Partial | ||
| 14 | xAI | 91.24 / 100 | Gratuit | Partial | ||
| 15 | Alibaba | 91.23 / 100 | Non disponible | Partial | ||
| 16 | Anthropic | 91.08 / 100 | 23 €/mois | Partial | ||
| 17 | Anthropic | 90.73 / 100 | 23 €/mois | Partial | ||
| 18 | xAI | 90.72 / 100 | Gratuit | Partial | ||
| 19 | xAI | 90.56 / 100 | Gratuit | Partial | ||
| 20 | 90.16 / 100 | 22 €/mois | Reliable | |||
| 21 | Baidu | 90.03 / 100 | Non disponible | Partial | ||
| 22 | Mimo V2.5 Pro | Xiaomi | 89.86 / 100 | Non disponible | Partial | |
| 23 | xAI | 89.69 / 100 | Gratuit | Partial | ||
| 24 | xAI | 89.51 / 100 | Gratuit | Partial | ||
| 25 | Alibaba | 89.50 / 100 | Non disponible | Partial | ||
| 26 | xAI | 89.37 / 100 | Gratuit | Reliable | ||
| 27 | DeepSeek | 89.34 / 100 | Gratuit | Partial | ||
| 28 | Anthropic | 88.99 / 100 | 23 €/mois | Partial | ||
| 29 | OpenAI | 88.86 / 100 | 23 €/mois | Reliable | ||
| 30 | xAI | 88.46 / 100 | Gratuit | Partial | ||
| 31 | 88.29 / 100 | 22 €/mois | Partial | |||
| 32 | Anthropic | 88.11 / 100 | 23 €/mois | Partial | ||
| 33 | ByteDance | 88.10 / 100 | Non disponible | Partial | ||
| 34 | DeepSeek | 88.03 / 100 | Gratuit | Reliable | ||
| 35 | Anthropic | 87.99 / 100 | 23 €/mois | Reliable | ||
| 36 | DeepSeek | 87.76 / 100 | Gratuit | Partial | ||
| 37 | xAI | 87.39 / 100 | Gratuit | Reliable | ||
| 38 | Anthropic | 86.89 / 100 | 23 €/mois | Partial | ||
| 39 | OpenAI | 86.88 / 100 | 23 €/mois | Partial | ||
| 40 | Moonshot AI | 86.87 / 100 | Non disponible | Partial | ||
| 41 | Baidu | 86.71 / 100 | Non disponible | Partial | ||
| 42 | Alibaba | 86.58 / 100 | Non disponible | Reliable | ||
| 43 | Z.ai (zhipu ai) | 86.52 / 100 | Non disponible | Reliable | ||
| 44 | Baidu | 86.36 / 100 | Non disponible | Partial | ||
| 45 | OpenAI | 86.35 / 100 | 23 €/mois | Partial | ||
| 46 | Anthropic | 86.23 / 100 | 23 €/mois | Partial | ||
| 47 | OpenAI | 85.66 / 100 | 23 €/mois | Partial | ||
| 48 | Inkling | Thinky | 85.49 / 100 | Non disponible | Partial | |
| 49 | OpenAI | 85.41 / 100 | 23 €/mois | Reliable | ||
| 50 | OpenAI | 85.32 / 100 | 23 €/mois | Reliable |
Score indicatif, pondéré à partir de sources publiques. Méthodologie transparente. Voir la page Classement IA pour la méthodologie complète et les autres catégories.
AI chat model ranking
Le Recul's Chat ranking compares the large language models people use every day to talk, write, summarise and reason: Claude, ChatGPT (GPT), Gemini, Grok, DeepSeek, Mistral, Qwen and their variants. Unlike marketing comparisons, it does not try to crown a single « best tool »: a model that excels at long-form reasoning can disappoint on factual accuracy, or cost too much for regular use.
The score aggregates several public signals — user preference, reasoning quality, factual accuracy, long-context handling, speed and subscription cost — weighted according to a documented « Chat » grid. The aim is to reflect what a model is worth in a real conversation, not how it performs on one isolated benchmark.
This ranking measures a trend, not a fixed truth. Models move fast, the weightings are public, and every evaluation carries a status (Reliable, Partial, Insufficient) depending on source coverage. Read it as a reference point for choosing according to your use — writing, support, analysis, light coding — not as a definitive league table.
What this ranking is for
This ranking helps you choose a conversational assistant for a specific use rather than on the hype of the moment. A model can dominate on reasoning and stay mediocre on factual accuracy, or be very fast but limited on long context.
It also helps put announcements in perspective: a new version presented as « revolutionary » does not necessarily change the ranking until its real gains are confirmed by public sources.
Comment lire le classement
- User preference: what humans actually prefer in blind tests, not launch-day scores.
- Reasoning and factual accuracy: following a complex instruction without making things up.
- Long context: holding up across long documents or conversations.
- Public monthly subscription cost, not the per-token API price.
- Evaluation status (Reliable / Partial / Insufficient) depending on source coverage.
FAQ — Chat
Which is the best chat model?
There is no universal « best ». The ranking reflects a balance between reasoning, factual accuracy, context, speed and cost. The right choice depends on your use: writing, customer support, analysis or light coding.
How often is the ranking updated?
The data is refreshed automatically about every 25 hours from public sources. The update date is shown at the top of the page.
Why is a well-known model marked « Partial »?
The status depends on the coverage of the available public sources. « Partial » means the data is not yet enough for a firm verdict, not that the model is bad.
Is the cost taken into account the API price?
No. For consumer chat, it is the public monthly subscription (Pro / Plus / Advanced) that is considered, not the API's per-token cost.
How is the score calculated?
It is a composite score weighted according to a documented « Chat » grid. The full methodology is available on the main ranking page.