AI ranking

Code

Solving real bugs (SWE-bench), file editing (Aider), long-form reasoning, reliability, cost.

Ranking updated on 02 October 2026 at 10:14

Top visuel

  1. 01 Deepseek r1 + claude 3 5 Sonnet 20241022 96.83
  2. 02 Claude Opus 5 5 Max 95.94
  3. 03 O3 (high) + gpt 4.1 95.80
  4. 04 GPT 6.1 sol (max) 90.72
  5. 05 Claude Sonnet 5.5 Xhigh 90.19
  6. 06 O3 (high) 85.83
  7. 07 Gemini 2.5 Pro Preview 06 05 (32k think) 85.04
  8. 08 Deepseek V3.2 Exp (reasoner) 84.39
  9. 09 Claude Fable 5 1 Max 84.07
  10. 10 Gemini 2.5 Pro Preview 06 05 (default think) 83.27

Comparateur Le Recul

Who wins on what

Cochez au moins deux modèles dans le tableau du classement, puis revenez ici.

# Model Vendor Score Monthly subscription Assessment
01 Anthropic 96.83 / 100 Gratuit Partial
02 Anthropic 95.94 / 100 23 €/mois Partial
03 OpenAI 95.80 / 100 23 €/mois Partial
04 OpenAI 90.72 / 100 23 €/mois Reliable
05 Anthropic 90.19 / 100 23 €/mois Partial
06 OpenAI 85.83 / 100 23 €/mois Partial
07 Google 85.04 / 100 22 €/mois Partial
08 DeepSeek 84.39 / 100 Gratuit Partial
09 Anthropic 84.07 / 100 23 €/mois Partial
10 Google 83.27 / 100 22 €/mois Partial
11 OpenAI 83.07 / 100 23 €/mois Reliable
12 Google 82.22 / 100 22 €/mois Partial
13 Anthropic 81.99 / 100 23 €/mois Partial
14 DeepSeek 81.79 / 100 Gratuit Partial
15 Google 81.69 / 100 22 €/mois Reliable
16 xAI 81.49 / 100 Gratuit Partial
17 Google 78.48 / 100 22 €/mois Partial
18 Spacexai 78.37 / 100 Gratuit Reliable
19 Anthropic 78.12 / 100 23 €/mois Partial
20 Mimo V2.6 Pro Xiaomi 77.37 / 100 Non disponible Reliable
21 DeepSeek 76.80 / 100 Gratuit Reliable
22 Anthropic 76.47 / 100 23 €/mois Partial
23 Anthropic 75.74 / 100 23 €/mois Partial
24 OpenAI 75.70 / 100 23 €/mois Partial
25 Anthropic 75.48 / 100 23 €/mois Partial
26 Anthropic 74.59 / 100 23 €/mois Partial
27 Anthropic 74.34 / 100 23 €/mois Partial
28 DeepSeek 73.90 / 100 Gratuit Partial
29 Anthropic 73.82 / 100 23 €/mois Partial
30 OpenAI 73.30 / 100 23 €/mois Reliable
31 Anthropic 72.84 / 100 23 €/mois Partial
32 Google 72.23 / 100 22 €/mois Partial
33 xAI 71.86 / 100 Gratuit Partial
34 OpenAI 71.34 / 100 23 €/mois Partial
35 Tencent 71.07 / 100 Non disponible Partial
36 Anthropic 70.94 / 100 23 €/mois Partial
37 xAI 70.62 / 100 Gratuit Partial
38 Anthropic 70.54 / 100 23 €/mois Partial
39 DeepSeek 69.93 / 100 Gratuit Partial
40 Anthropic 69.50 / 100 23 €/mois Partial
41 Alibaba 69.12 / 100 Non disponible Partial
42 OpenAI 68.80 / 100 23 €/mois Partial
43 OpenAI 67.43 / 100 23 €/mois Reliable
44 Quasar alpha — 67.13 / 100 Non disponible Partial
45 Ling 3.0 Flash Fin Inclusionai 66.69 / 100 Non disponible Partial
46 Google 66.18 / 100 22 €/mois Partial
47 Google 65.94 / 100 22 €/mois Partial
48 Optimus alpha — 65.59 / 100 Non disponible Partial
49 Anthropic 65.35 / 100 23 €/mois Partial
50 Alibaba 65.03 / 100 Non disponible Partial

Score indicatif, pondéré à partir de sources publiques. Méthodologie transparente. Voir la page Classement IA pour la méthodologie complète et les autres catégories.

AI coding model ranking

Le Recul's Code ranking assesses AI models on what actually matters for development: fixing real bugs, editing existing files, holding a long line of reasoning across a codebase and staying reliable from one run to the next. It draws on public signals such as SWE-bench (solving real tickets) and Aider (code editing), weighted according to a « Code » grid.

Here too, there is no « best tool » in the abstract. A model can shine on isolated exercises and fail on a real codebase, or perform well but cost too much for intensive use. The ranking separates code models (Claude, GPT, Gemini, DeepSeek, Qwen Coder…) from agentic development systems (Claude Code, Cursor, Aider), which are handled separately under Coding agents.

The score reflects a trend measured on real tasks, not a commercial promise. Every evaluation carries a reliability status depending on source coverage, and the weightings stay public. Use it to choose a model for your context — language, project size, budget — rather than to follow the latest announcement.

What this ranking is for

This ranking helps you choose a model to assist development: completion, refactoring, bug fixing, test generation. It deliberately separates raw models from coding agents, which add a layer of tools and autonomy.

It also helps temper spectacular demos: succeeding on a scripted exercise does not guarantee reliability on a real repository, with its dependencies and its history.

Comment lire le classement

  • Solving real bugs (SWE-bench style), not toy exercises.
  • Ability to edit existing files (Aider style).
  • Long-form reasoning across a codebase, without losing the thread.
  • Reliability and reproducibility from one run to the next.
  • Cost relative to how intensively you use it.

FAQ — Code

What is the difference between a code model and a coding agent?

A code model generates and edits code as text. A coding agent (Claude Code, Cursor, Aider) adds tools, runs, tests and fixes in a loop. Agents are ranked separately under Coding agents.

What is the Code score based on?

Mainly on solving real bugs (SWE-bench) and code editing (Aider), completed by reasoning, reliability and cost, according to a documented « Code » grid.

Will a highly ranked model be reliable on my project?

The ranking gives a general trend. Your results depend on the language, the size of the repository and the context. Test on a representative task before concluding.

Are open-weights models included?

Yes, where the sources allow it. Open models that run locally also have a dedicated view (Local AI — code).

How often is this updated?

About every 25 hours, automatically, from public sources. The date is shown at the top of the page.

See the full FAQ →

Related rankings