AI ranking
Code
Solving real bugs (SWE-bench), file editing (Aider), long-form reasoning, reliability, cost.
Top visuel
- 01 Deepseek r1 + claude 3 5 Sonnet 20241022 96.83
- 02 Claude Opus 5 5 Max 95.94
- 03 O3 (high) + gpt 4.1 95.80
- 04 GPT 6.1 sol (max) 90.72
- 05 Claude Sonnet 5.5 Xhigh 90.19
- 06 O3 (high) 85.83
- 07 Gemini 2.5 Pro Preview 06 05 (32k think) 85.04
- 08 Deepseek V3.2 Exp (reasoner) 84.39
- 09 Claude Fable 5 1 Max 84.07
- 10 Gemini 2.5 Pro Preview 06 05 (default think) 83.27
Compare models. Cochez jusqu'à 3 modèles dans le tableau ci-dessous, puis lancez la comparaison : qui gagne sur quel usage, de combien, et avec quelle solidité.
| # | Model | Vendor | Score | Monthly subscription | Assessment | |
|---|---|---|---|---|---|---|
| 01 | Anthropic | 96.83 / 100 | Gratuit | Partial | ||
| 02 | Anthropic | 95.94 / 100 | 23 €/mois | Partial | ||
| 03 | OpenAI | 95.80 / 100 | 23 €/mois | Partial | ||
| 04 | OpenAI | 90.72 / 100 | 23 €/mois | Reliable | ||
| 05 | Anthropic | 90.19 / 100 | 23 €/mois | Partial | ||
| 06 | OpenAI | 85.83 / 100 | 23 €/mois | Partial | ||
| 07 | 85.04 / 100 | 22 €/mois | Partial | |||
| 08 | DeepSeek | 84.39 / 100 | Gratuit | Partial | ||
| 09 | Anthropic | 84.07 / 100 | 23 €/mois | Partial | ||
| 10 | 83.27 / 100 | 22 €/mois | Partial | |||
| 11 | OpenAI | 83.07 / 100 | 23 €/mois | Reliable | ||
| 12 | 82.22 / 100 | 22 €/mois | Partial | |||
| 13 | Anthropic | 81.99 / 100 | 23 €/mois | Partial | ||
| 14 | DeepSeek | 81.79 / 100 | Gratuit | Partial | ||
| 15 | 81.69 / 100 | 22 €/mois | Reliable | |||
| 16 | xAI | 81.49 / 100 | Gratuit | Partial | ||
| 17 | 78.48 / 100 | 22 €/mois | Partial | |||
| 18 | Spacexai | 78.37 / 100 | Gratuit | Reliable | ||
| 19 | Anthropic | 78.12 / 100 | 23 €/mois | Partial | ||
| 20 | Mimo V2.6 Pro | Xiaomi | 77.37 / 100 | Non disponible | Reliable | |
| 21 | DeepSeek | 76.80 / 100 | Gratuit | Reliable | ||
| 22 | Anthropic | 76.47 / 100 | 23 €/mois | Partial | ||
| 23 | Anthropic | 75.74 / 100 | 23 €/mois | Partial | ||
| 24 | OpenAI | 75.70 / 100 | 23 €/mois | Partial | ||
| 25 | Anthropic | 75.48 / 100 | 23 €/mois | Partial | ||
| 26 | Anthropic | 74.59 / 100 | 23 €/mois | Partial | ||
| 27 | Anthropic | 74.34 / 100 | 23 €/mois | Partial | ||
| 28 | DeepSeek | 73.90 / 100 | Gratuit | Partial | ||
| 29 | Anthropic | 73.82 / 100 | 23 €/mois | Partial | ||
| 30 | OpenAI | 73.30 / 100 | 23 €/mois | Reliable | ||
| 31 | Anthropic | 72.84 / 100 | 23 €/mois | Partial | ||
| 32 | 72.23 / 100 | 22 €/mois | Partial | |||
| 33 | xAI | 71.86 / 100 | Gratuit | Partial | ||
| 34 | OpenAI | 71.34 / 100 | 23 €/mois | Partial | ||
| 35 | Tencent | 71.07 / 100 | Non disponible | Partial | ||
| 36 | Anthropic | 70.94 / 100 | 23 €/mois | Partial | ||
| 37 | xAI | 70.62 / 100 | Gratuit | Partial | ||
| 38 | Anthropic | 70.54 / 100 | 23 €/mois | Partial | ||
| 39 | DeepSeek | 69.93 / 100 | Gratuit | Partial | ||
| 40 | Anthropic | 69.50 / 100 | 23 €/mois | Partial | ||
| 41 | Alibaba | 69.12 / 100 | Non disponible | Partial | ||
| 42 | OpenAI | 68.80 / 100 | 23 €/mois | Partial | ||
| 43 | OpenAI | 67.43 / 100 | 23 €/mois | Reliable | ||
| 44 | Quasar alpha | — | 67.13 / 100 | Non disponible | Partial | |
| 45 | Ling 3.0 Flash Fin | Inclusionai | 66.69 / 100 | Non disponible | Partial | |
| 46 | 66.18 / 100 | 22 €/mois | Partial | |||
| 47 | 65.94 / 100 | 22 €/mois | Partial | |||
| 48 | Optimus alpha | — | 65.59 / 100 | Non disponible | Partial | |
| 49 | Anthropic | 65.35 / 100 | 23 €/mois | Partial | ||
| 50 | Alibaba | 65.03 / 100 | Non disponible | Partial |
Score indicatif, pondéré à partir de sources publiques. Méthodologie transparente. Voir la page Classement IA pour la méthodologie complète et les autres catégories.
AI coding model ranking
Le Recul's Code ranking assesses AI models on what actually matters for development: fixing real bugs, editing existing files, holding a long line of reasoning across a codebase and staying reliable from one run to the next. It draws on public signals such as SWE-bench (solving real tickets) and Aider (code editing), weighted according to a « Code » grid.
Here too, there is no « best tool » in the abstract. A model can shine on isolated exercises and fail on a real codebase, or perform well but cost too much for intensive use. The ranking separates code models (Claude, GPT, Gemini, DeepSeek, Qwen Coder…) from agentic development systems (Claude Code, Cursor, Aider), which are handled separately under Coding agents.
The score reflects a trend measured on real tasks, not a commercial promise. Every evaluation carries a reliability status depending on source coverage, and the weightings stay public. Use it to choose a model for your context — language, project size, budget — rather than to follow the latest announcement.
What this ranking is for
This ranking helps you choose a model to assist development: completion, refactoring, bug fixing, test generation. It deliberately separates raw models from coding agents, which add a layer of tools and autonomy.
It also helps temper spectacular demos: succeeding on a scripted exercise does not guarantee reliability on a real repository, with its dependencies and its history.
Comment lire le classement
- Solving real bugs (SWE-bench style), not toy exercises.
- Ability to edit existing files (Aider style).
- Long-form reasoning across a codebase, without losing the thread.
- Reliability and reproducibility from one run to the next.
- Cost relative to how intensively you use it.
FAQ — Code
What is the difference between a code model and a coding agent?
A code model generates and edits code as text. A coding agent (Claude Code, Cursor, Aider) adds tools, runs, tests and fixes in a loop. Agents are ranked separately under Coding agents.
What is the Code score based on?
Mainly on solving real bugs (SWE-bench) and code editing (Aider), completed by reasoning, reliability and cost, according to a documented « Code » grid.
Will a highly ranked model be reliable on my project?
The ranking gives a general trend. Your results depend on the language, the size of the repository and the context. Test on a representative task before concluding.
Are open-weights models included?
Yes, where the sources allow it. Open models that run locally also have a dedicated view (Local AI — code).
How often is this updated?
About every 25 hours, automatically, from public sources. The date is shown at the top of the page.