Code
See the ranking →- 01 96.8 Anthropic
- 02 95.9 Anthropic
- 03 95.8 OpenAI
Section
Comparing AI without the hype. General-purpose LLMs, coding agents, browser agents, business agents, open-source models. An indicative score, weighted from public sources.
Compare models. Cochez jusqu'à 3 modèles dans le tableau ci-dessous, puis lancez la comparaison : qui gagne sur quel usage, de combien, et avec quelle solidité.
Les modèles au-delà d'environ 35 milliards de paramètres ne sont pas classés dans les catégories locales. Le Recul privilégie les modèles réellement exploitables en local, sans infrastructure professionnelle lourde.
| # | Model | Vendor | Le Recul score | Monthly subscription | Assessment | |
|---|---|---|---|---|---|---|
| 01 | Anthropic | 96.50 / 100 | 23 €/mois | Partial | ||
| 02 | Anthropic | 96.33 / 100 | 23 €/mois | Partial | ||
| 03 | Anthropic | 95.98 / 100 | 23 €/mois | Partial | ||
| 04 | Anthropic | 95.10 / 100 | 23 €/mois | Partial | ||
| 05 | Anthropic | 94.71 / 100 | 23 €/mois | Reliable | ||
| 06 | Anthropic | 94.58 / 100 | 23 €/mois | Partial | ||
| 07 | Anthropic | 94.06 / 100 | 23 €/mois | Partial | ||
| 08 | Anthropic | 92.31 / 100 | 23 €/mois | Partial | ||
| 09 | Anthropic | 92.23 / 100 | 23 €/mois | Reliable | ||
| 10 | 92.07 / 100 | 22 €/mois | Reliable | |||
| 11 | OpenAI | 91.43 / 100 | 23 €/mois | Partial | ||
| 12 | Anthropic | 91.26 / 100 | 23 €/mois | Partial | ||
| 13 | OpenAI | 91.25 / 100 | 23 €/mois | Partial | ||
| 14 | xAI | 91.24 / 100 | Gratuit | Partial | ||
| 15 | Alibaba | 91.23 / 100 | Non disponible | Partial | ||
| 16 | Anthropic | 91.08 / 100 | 23 €/mois | Partial | ||
| 17 | Anthropic | 90.73 / 100 | 23 €/mois | Partial | ||
| 18 | xAI | 90.72 / 100 | Gratuit | Partial | ||
| 19 | xAI | 90.56 / 100 | Gratuit | Partial | ||
| 20 | 90.16 / 100 | 22 €/mois | Reliable |
Data temporarily unavailable. The ranking is updated once a day by Le Recul's BENCH pipeline. If nothing is showing, try again in a few moments.
Recent models, previews or challengers with a promising signal, but not yet documented well enough to enter a main ranking.
The ranking draws on several public benchmarks to compare models by use: conversation, code, agents, cost and performance. The Le Recul score is the final composite score calculated from those sources — not a source added to the calculation.
Compares models on public multi-format evaluations (text, web code, image, video) through user votes and a Bradley-Terry calculation.
Current best model: Claude Opus 4 6 High · 96.5
An index of model quality, speed and latency via the official Artificial Analysis API (awaiting a free API key).
Current best model: Gemini 4 argon (high) · 83.9
Measures the ability of agents and scaffolds to fix real bugs drawn from open source projects.
A multi-language variant (Python/Java/TS/Go/Rust/C#/C++/Windows) with verified instances renewed every month.
Measures performance on multi-language coding tasks. The classic modes feed Code, architect mode feeds Coding agents.
Current best model: Grok 3 mini beta (high) · 89.4
Feeds the internal knowledge criterion of the Le Recul score: 14 disciplines (biology, law, maths, physics, philosophy…) via TIGER-Lab's official dataset. Not displayed as a public column.
Current best model: Gemini 3 Pro · 92.1
The reasoning component of the Le Recul score: a hard extension of BIG-Bench, DeepMind's official leaderboard.
Current best model: Deepseek r1 · 77.8
Current best model: Claude Opus 4 8 Xhigh Effort · 79.2
Current best model: Claude Fable 5 High · 94.7
The Le Recul score is indicative and composite: it aggregates several public sources (LM Arena, SWE-bench, Aider Polyglot, and Artificial Analysis once the API is connected), normalises the values on a 0–100 scale and applies a documented weighting per category. This score is not a source added to the calculation — it is the result of the aggregation. No data is invented: if a category lacks measurements, the display says “insufficient data”.
The Assessment column reflects source coverage for each model: Reliable when at least two relevant independent sources confirm the model (cross-validation), or when a strong reference source (LM Arena, Artificial Analysis, LiveBench, Epoch) on its own covers at least 60 % of the category's weighted criteria — preview, beta and experimental models, along with effort or reasoning variants, are excluded from that promotion and stay “Partial”; Partial when a single source documents the model with usable but limited coverage — often a recent or single-source model, not yet cross-checked by a second independent source; Insufficient when the data does not allow a clean verdict. A knowledge measurement alone (MMLU-Pro) is never enough to make a model “Reliable”, and categories where no strong source covers enough criteria (for example Code) stay “Partial” until a second independent source is available.
Le Recul score. The score shown is the settled Le Recul score: it combines the scores of reliable public sources (LM Arena, SWE-bench, Aider, MMLU-Pro, BBEH, LiveBench, Artificial Analysis once the API is connected) according to a documented weighting per category. Internally, a tiny tie-breaking component (≤ 0.01) folds in coverage, the number of independent sources, the assessment status and how recent the data is — invisible at two decimal places for normal scores, it separates models that look equivalent. The score shown is therefore a Le Recul score, not a raw average: apparent ties are settled by coverage, stability and relevance of the sources. The score shown is limited to two decimal places. No model is removed or hidden in the event of a tie.
The Cost / performance category compares overall performance with the price of the public monthly subscription (Claude.ai Pro, ChatGPT Plus, Gemini Advanced, Le Chat Pro, free interfaces…). API prices and per-token costs are not taken into account, because they do not reflect what an ordinary user actually pays.
The Local / open-source AI category gathers open-weights or open-source models that can be used locally without heavy infrastructure. Very large models (beyond roughly 35 billion parameters) are excluded from the main local ranking by editorial choice. Proprietary cloud models (Claude, GPT, Gemini, Grok, Mistral cloud…) never appear in this category.
The Coding agents category gathers agentic systems (scaffolds + models) measured on complete coding tasks: SWE-bench (solving real bugs) and Aider in architect mode. Models on their own (Aider whole/diff) stay in the Code category.
Additional components. BBEH (DeepMind's official leaderboard) feeds the reasoning criterion; MMLU-Pro (TIGER-Lab's official dataset) feeds an internal knowledge criterion (14 disciplines). Both benchmarks count as independent sources for the Le Recul score — never displayed as a separate public column, and never counted twice. Their effect stays measured: they mainly help to cross-check models already covered by other sources; a recent model that MMLU-Pro does not yet cover stays “Partial”. Benchmarks for which no usable source has been identified (IFEval, classic BBH, MATH, GPQA, MuSR) are switched on as soon as the Artificial Analysis connection is available.
Hidden categories. Some planned categories (agent families in particular) may be temporarily hidden from the public ranking as long as no source provides enough rankable models. They are not deleted: they reappear automatically as soon as a collection run provides sufficient data. A category that contains models is never hidden, even if many of its assessments are still “Partial”.
Data continuity. If a source does not respond during a cycle, the ranking is not emptied: that source's last known reliable measurement is kept, within a freshness limit. A guardrail blocks any publication that would abruptly degrade the ranking (a critical source lost, an abnormal drop). No data that is too old is presented as fresh.
An indicative score, weighted from public sources. Transparent methodology.
Cadence : Mise à jour 1×/jour, heure variable. The site is not rebuilt on every cycle: only the data files are updated.
Each category aggregates recognised public sources (LM Arena, SWE-bench, Aider, MMLU-Pro, LiveBench, Artificial Analysis and others), weighted by use case. Le Recul derives a composite score normalised to 100, with close calls settled on source coverage, source independence, evaluation status and recency. The detailed method is published on the AI ranking page, in French.
A model may be missing when public data is lacking, rests on a single uncorroborated source, or does not fit the category — a local model too large for the « local » categories, for instance. Absence is not a judgement: it means the data is too thin for an honest ranking.
The data is refreshed regularly, and the cadence is shown at the top of the ranking. Only the data files are updated, not the whole site, and the date of the last update is displayed at the top of the page.
An AI that is excellent at code is not necessarily good at images, chat or agent work. Ranking by use case avoids the misleading comparisons of a generic « top 10 » and reflects what choosing a tool actually involves.
No. The ranking is neither sponsored nor paid for: no vendor can buy a position. Affiliate links, where they exist, are stated and play no part in the scores.