On 10 September 2026, DeepSeek released V4.1 Flash, the first model in its V4.1 generation: 552 billion parameters, open weights under the MIT licence, available the same day on its website, in its app and through its API. The number everyone repeated comes in three parts: $0.003 per million tokens re-read from cache, $0.15 on input, $0.60 on output.

Update, 26 September 2026. We wrote that the coding score would need a few days to be verified. The official DeepSWE leaderboard has indeed been refreshed since, on 22 September — without adding V4.1 Flash. It is not measured on compar:IA either. The 74.2% remains a vendor figure, sixteen days after release. The compar:IA numbers have also moved; they are read again further down.

In its own model card, DeepSeek measures itself against Claude Opus 5, GPT-5.6 Sol, Kimi K3 and GLM-5.3 — the first two of which sell for more than ten times the price. It never compares itself to the model it actually competes with on price: GPT-5.6 Luna, the entry point of OpenAI’s GPT-5.6 family, billed at $0.20 on input and $1.20 on output, and the engine behind free ChatGPT since August.

That is the matchup this comparison measured. And the result turns less on the scores than on a question nobody asks: what time do you call the model?


The verdict in brief

On intelligence: DeepSeek, narrowly. On version 4.3 of the Artificial Analysis index, V4.1 Flash scores 40, GPT-5.6 Luna 38 at its maximum setting. Across the ten evaluations, DeepSeek wins five, Luna three, two are ties.

On price: it depends on the clock, and on your continent. The advertised rate is the off-peak rate. Monday to Friday, 01:00–04:00 and 06:00–10:00 UTC, DeepSeek charges double: $0.30 on input — more than Luna — and $1.20 on output, exactly Luna’s price. Those windows were drawn around the Beijing working day, which produces a result most buyers miss: an American working day never touches them.

On the actual bill: Luna, most of the time. DeepSeek writes 2.2 times more tokens than Luna to do the same work. Running the full index costs $477 at DeepSeek’s peak rates against $320 at OpenAI. At off-peak rates DeepSeek goes back in front, at roughly $238.

On code: DeepSeek where it is measured, a promise everywhere else. DeepSeek claims 74.2% on DeepSWE, which would put it top of the official leaderboard; that score still does not appear there, even though the leaderboard was refreshed on 22 September 2026. The independent measurement that does exist, on command-line work, is unambiguous: 27% against 12%.

On speed: DeepSeek, by a distance. A complete answer lands in 12.6 seconds, against 131 for Luna at its maximum setting, reasoning included.

On free access and data: two opposite logics. Free ChatGPT works without even an account and gives unlimited text to free accounts. DeepSeek requires sign-up and stores your data in China. Both may use your conversations to improve their models. But only DeepSeek offers a fire exit: its open weights, which other hosts serve — or which you can install yourself.


Where these numbers come from

Method before figures, because a comparison is worth only what its sourcing is worth. Four families of data, kept separate throughout:

  • Vendor figures come from DeepSeek’s and OpenAI’s announcements, price lists and model cards. They are authoritative on prices, capabilities and limits. They are suspect on comparisons: each vendor picks its own evaluations and its own opponents, and DeepSeek, as we saw, never lines itself up against Luna.
  • Independent measurements come from Artificial Analysis, which runs the same battery on both models and publishes what the vendors leave out — tokens consumed, real cost, latency — and from the official DeepSWE leaderboard, run by Datacurve, which measures coding level by level.
  • User preferences come from compar:IA, the public comparator run by France’s Ministry of Culture, where visitors blind-judge two models’ answers. Its voters are French-speaking, which is a limit we state rather than hide: it is the only open preference data available on these two models.
  • Our own ranks come from Le Recul’s AI ranking, read on 16 September 2026. We only place entries side by side when they are fed by the same family of sources: two scores from two different harnesses are not comparable.

Every cost calculation was redone from the published price lists, with the assumptions written out so you can replay them.

Two gaps are flagged rather than filled: the quality of writing outside English, which neither vendor evaluates publicly, and V4.1 Flash’s coding score, which the official leaderboard still has not measured — the model shipped on 10 September, and the leaderboard was refreshed on the 22nd without it.


The two contenders

  DeepSeek V4.1 Flash GPT-5.6 Luna
VendorDeepSeek (China)OpenAI (United States)
Released10 September 20269 July 2026; the free ChatGPT model since August
API identifierdeepseek-flashgpt-5.6-luna
Architecturemixture of experts, 552 billion parameters, 8 billion active on read and 16 on writenot stated on the model page
Weightsopen, MIT licenceclosed
Context1,000,000 tokens1,050,000 tokens, of which 922,000 maximum on input
Maximum output384,000 tokens128,000 tokens
Input / outputtext and images / texttext and images / text
Reasoningon by default, effort adjustable from 1 to 100none, low, medium (default), high, xhigh, max
Price per million tokens$0.30 / $1.20 at peak; $0.15 / $0.60 off-peak$0.20 / $1.20 at any hour
Cached input$0.006 (peak); $0.003 (off-peak)$0.02
Long-context surchargenone in the price list×2 on input and ×1.5 on output beyond 272,000 tokens

Two rows in this table matter more than they look. At DeepSeek, reasoning is on by default, and the scores on its card are obtained at maximum effort, 100. At OpenAI, the API’s default setting is medium — and we will see that it changes everything.


The advertised price is the off-peak price — and it is billed in UTC

DeepSeek bills by the clock. Its price list distinguishes peak hours, Monday to Friday from 01:00 to 04:00 and from 06:00 to 10:00 UTC, from off-peak hours — everything else in the week, weekends included — billed at exactly half price.

Those windows were drawn around China: in Beijing they cover 09:00–12:00 and 14:00–18:00, the working day. Which leads to the finding this comparison did not expect.

A 9am–6pm working day in… Peak hours hit, out of 9 Share of the day at double price
New York (EDT, and EST from 1 November)0none
San Francisco (PDT, to 31 October)0none
San Francisco (PST, from 1 November)111 %
London (BST, to 24 October)222 %
London (GMT, from 25 October)111 %
Paris, Berlin (CEST, to 24 October)333 %
Paris, Berlin (CET, from 25 October)222 %

For an American team, DeepSeek’s peak-hour trap does not exist. Both windows fall between late evening and dawn on the East Coast, and outside office hours on the West Coast — one hour a day at most, once the clocks change in November. An American working day is billed off-peak from start to finish, which means the $0.15 everyone quotes is, for once, the price you actually pay.

For a European team, the same windows land squarely on the morning. That is the whole difference, and it comes from nothing but geography.

And here is what the comparison becomes, line by line:

Per million tokens DeepSeek, peak DeepSeek, off-peak GPT-5.6 Luna
Cached input$0.006$0.003$0.02
Standard input$0.30$0.15$0.20
Output$1.20$0.60$1.20

Off-peak, DeepSeek is 25% cheaper than Luna on input and twice as cheap on output. At peak it is 50% more expensive on input and the same price on output. Only the cache stays clearly in its favour at any hour, three to seven times cheaper.


The token price is not the bill

A rate per million tokens says nothing about how many tokens a model burns to do its job. That is precisely where the two contenders diverge most.

Artificial Analysis publishes, for each model, what it cost to run its entire index — the same ten evaluations, at the vendor’s rates, counting cache, reasoning and answer. It also publishes the number of tokens written per task.

Measured on the same battery DeepSeek V4.1 Flash (max) Luna (max) Luna (xhigh)
Intelligence index v4.3403835
Tokens written per task89,00041,00024,000
of which reasoning63,00028,00015,000
Measured cost of the full index (DeepSeek at peak rates)$477$320$179
Same cost, DeepSeek at off-peak rates (our calculation)about $238$320$179

DeepSeek writes 2.2 times more than Luna at its maximum setting, and 3.7 times more than Luna one notch below. Its reasoning alone outweighs Luna’s entire output at maximum.

Hence a reversal the price list does not show: at peak rates — the ones Artificial Analysis used — DeepSeek costs 49% more than Luna to sit the same exam. At off-peak rates it costs 25% less. The last row is our own calculation, and it is exact: since every DeepSeek rate is halved, so is the bill.

Luna’s xhigh setting deserves a detour: three index points below maximum, for a bill 1.8 times lighter. At that setting Luna stays cheaper than DeepSeek at any hour — but it becomes the least intelligent of the three. We met the same arbitration between price and reliability in our comparison of Kimi K3 against Claude Fable 5, where the cheapest model solved more tasks per dollar while hallucinating far more.

DeepSeek is open about the mechanism. According to figures relayed by VentureBeat, raising its reasoning effort from 25 to 100 lifts its coding score from 66.0% to 74.2%, for roughly 2.5 times more output tokens. Quality is bought in volume. The advertised price is the price of the token; the bill is the price of the work. Our cost / performance ranking tracks this trade-off across the whole market.


The calculation that decides: how many tokens before DeepSeek costs more

Take a common use — an assistant embedded in a business tool — and write the assumptions so anyone can replay them:

  • 100,000 requests a month;
  • 3,000 input tokens per request, of which 2,000 are repeated instructions re-read from cache and 1,000 are new;
  • 500 output tokens per request with Luna.

At OpenAI, the bill is $84 a month, at any hour.

At DeepSeek it depends on two unknowns: the hour, and how many tokens the model will write to do the same work. Call that ratio k: k = 1 if it writes as much as Luna, k = 2 if it writes twice as much. The ratio measured by Artificial Analysis on its index is 2.17.

Monthly DeepSeek bill if k = 1 if k = 2.17 More expensive than Luna once k exceeds
Everything at peak$91.20$161.400.88
Working day in New York (0 peak hours)$45.60$80.702.28
Working day in San Francisco, winter, or London from 25 October (1 peak hour)$50.67$89.672.00
Working day in London, summer (2 peak hours)$55.73$98.631.77
Working day in Paris or Berlin, summer (3 peak hours)$60.80$107.601.58

Three readings. At full peak rates, DeepSeek costs more than Luna as soon as it writes 88% of what Luna writes — in other words, almost always. On an American working day, it stays cheaper until it writes 2.28 times more: the measured ratio, 2.17, slips just under the bar, so DeepSeek wins on cost, by a margin of about 4%. On a European working day the threshold drops to 1.77 or 1.58, and at the measured ratio DeepSeek costs 17% to 28% more than Luna.

The same model, the same rates, the same workload — and the cheaper option flips depending on which side of the Atlantic the calls come from. That is not a detail of the small print; on this scenario it is the entire decision.

One caveat that matters: the 2.17 ratio was measured on demanding evaluations, with reasoning pushed to maximum. On simple requests the gap can be very different. That is the one measurement to make on your own workload before choosing; the arithmetic itself stays the same.

And there is a third route that DeepSeek’s price list does not show. Because its weights are open, other hosts sell it on their own terms: Artificial Analysis lists seven providers, DeepSeek included, with prices varying by up to 4.9 times. Databricks lists it at $0.14 input and $0.28 output — cheaper than DeepSeek itself off-peak. Our ranking of open models lists the alternatives against the same criteria.

Real cost comparison between DeepSeek V4.1 Flash and GPT-5.6 Luna by time of use and token volume.

The advertised price per million tokens is not enough: depending on when you call and how much is actually generated, the final bill moves a long way.


Intelligence: 40 against 38, and a split that says everything

All the figures below come from version 4.3 of the Artificial Analysis index. That composite is versioned, and changing version changes the scores: values published under an earlier version, including in older comparisons, are not comparable to these.

Evaluation What it measures DeepSeek Luna (max)
AA-Briefcasecasework14241333
GDPval-AA v2knowledge work16321462
AutomationBench-AAautomated business workflows69 %50 %
Terminal-Bench 4.0command-line work27 %12 %
SciCodescientific code52 %54 %
Humanity's Last Examvery hard expert questions39 %39 %
GDP.pdfdocument reading13 %24 %
CritPtresearch physics14 %21 %
AA-Omnisciencereliability, hallucination−5−10
AA-LCR v1.1long-context reasoning84 %84 %

Five evaluations to DeepSeek, three to Luna, two ties. The two-point overall gap is small; how it is distributed is not.

DeepSeek wins wherever the job is to act: nineteen points ahead on business-workflow automation, more than double the score on the terminal, and a clear lead on casework. Luna wins wherever the job is to read: a PDF document, a research-physics problem.

That is the surprise in the table. DeepSeek promotes its native visual understanding, pre-trained on a corpus mixing images and text, and its card claims 95.6 on DocVQA for the base model. Yet document reading is where it gives way most clearly: 13% against 24%, almost half.

At the entry level, as at the top of the market, the decision is now made two index points at a time. It is no longer the total that should guide the choice, it is the row in the table that matches your use. Our general AI ranking is updated continuously and lets you place both models against the rest.


Code: a promise the official leaderboard has not verified

On coding, DeepSeek hits hard on its card: 74.2% on DeepSWE v1.1, a battery of 113 software-development tasks. That score would put it top of the official leaderboard run by Datacurve, ahead of GPT-6 Astra (74.1%), Gemini 3.8 Flash (73.8%) and Claude Opus 5 (73.6%).

The problem: as of 16 September 2026, V4.1 Flash is not on it. On the DeepSeek side, the official leaderboard stops at V4 Flash (53.3%) and V4 Pro (62.8%). The 74.2% therefore remains, to this day, a vendor figure.

And it has stayed one. We rechecked on 26 September 2026: Datacurve refreshed its leaderboard on 22 September, twelve days after the model shipped, and V4.1 Flash was not added. The page still stops at V4 Flash and V4 Pro. This is not a rebuttal — a battery of 113 development tasks is expensive to run, and Datacurve chooses its entries. But it does mean that sixteen days in, the only coding score available for V4.1 Flash is still the one its vendor publishes on its own card.

Should you distrust it? History rather argues for DeepSeek. For its two previous models, its card claimed 54.4% and 62.7%; the official leaderboard measured 53.3% and 62.8%. One point out on one, a tenth on the other. The promise is credible; it is not verified.

The one independent measurement already published on V4.1 Flash points the same way. On Terminal-Bench 4.0, Artificial Analysis records 27% against 12% for Luna at maximum. DeepSeek itself claimed 31.2%: the usual gap between a vendor’s harness and an independent evaluator’s.

Our AI ranking for code confirms the trend on a comparable basis. In the series fed by LMArena’s web-development arena, the only one where both models are measured the same way, V4.1 Flash ranks 27th and Luna, at its xhigh setting, 78th. For the specialised tools above them both, see our comparison of the best AI for coding.


The default-setting trap at OpenAI

The official DeepSWE leaderboard holds a piece of information nobody highlights: it measured Luna at every one of its reasoning levels. The drop is vertiginous.

GPT-5.6 Luna, setting Tasks solved first try Average cost per task
max67.2 %$0.61
xhigh56.9 %$0.31
high44.2 %$0.16
medium, the API default11.3 %$0.04
low1.5 %$0.01

At the default setting, Luna solves one coding task in nine. At maximum, two in three. Six times the success rate for fourteen times the spend — which, in absolute terms, is 61 cents per task.

This is the most expensive mistake in this comparison, and it is free to avoid: a developer who wires up Luna without touching the reasoning parameter is testing a model that bears almost no relation to the one in the leaderboards. The same vigilance applies at DeepSeek, whose reasoning is on by default, but whose published scores are obtained at maximum effort.


Speed: ten times faster, and why to be careful with that

Artificial Analysis measures writing at 218 tokens per second for DeepSeek against 115 for Luna at maximum. But the figure that matters to a user is the time to a complete 500-token answer, reasoning included: 12.6 seconds for DeepSeek, 131 seconds for Luna at maximum, 45.8 seconds at the xhigh setting.

More than two minutes of waiting for one answer: at its maximum setting Luna is not a conversational tool, it is a background-processing tool.

The usual caveat counts more here than elsewhere: speed is the least reliable number in any comparison. It depends on thinking time, which varies with the difficulty of the question, and even more on the host. For the same DeepSeek V4.1 Flash, Artificial Analysis records throughput ranging from 106 to 539 tokens per second depending on the provider.


The facts: which one makes up less?

AA-Omniscience, the Artificial Analysis reliability evaluation, rewards correct answers, penalises invented ones and does not punish abstention. Its scale runs from −100 to 100; zero means a model that is right as often as it is wrong.

Both are below zero: DeepSeek at −5, Luna at −10 at maximum and −11 at xhigh. On the knowledge questions in this evaluation, both models are therefore wrong slightly more often than they are right; DeepSeek, slightly less so than Luna.

OpenAI puts forward progress of its own: in an internal evaluation covering financial, medical and legal questions, Luna’s answers containing at least one factual error are said to be about 62% less frequent than with GPT-5.5 Instant. That is a vendor measurement, and the announcement does not detail the protocol.

On DeepSeek’s side, one figure from its own card deserves attention. On SimpleQA-Verified, which probes factual knowledge, the base model of V4.1 Flash scores 42.3, against 55.2 for V4 Pro’s. Which leads to the point DeepSeek does not highlight.


What DeepSeek does not highlight: its Pro model was replaced by the Flash

Since 14 September 2026 at 04:00 UTC, every request sent to deepseek-v4-pro is served by V4.1 Flash, at V4.1 Flash’s price, pending a V4.1 Pro with no announced date.

DeepSeek presents the switch as an upgrade, and for agentic work it is one: its card gives the new model 31.2% on Terminal-Bench 4.0 against 12.4% for the old Pro, and 74.2% against 62.7% on DeepSWE.

But the same card shows that V4.1 Flash knows less. For the base models: 42.3 against 55.2 on SimpleQA-Verified, 45.2 against 51.5 on LongBench-V2, the long-document comprehension evaluation, and 45.5 against 50.9 on MultiLoKo, which tests local knowledge across several languages. In their final versions, at maximum effort: 36.8 against 42.7 on Humanity’s Last Exam.

For an application that called V4 Pro in order to use its knowledge, the answer changed overnight, with no change of identifier. VentureBeat notes that migrating a model behind an existing identifier can invalidate regression tests. And DeepSeek itself concedes, according to the same source, that selection errors in its compressed-attention mechanism could degrade capability “in untested edge cases”.

The bill, on the other hand, collapsed: users of the old Pro now pay 4.4 times less on input and 3.3 times less on output, at peak rates. Four months after DeepSeek made a 75% price cut on V4 Pro permanent, that model is no longer served at all.


Long context: a threshold at OpenAI, none in DeepSeek’s price list

Both models accept about a million tokens. But OpenAI applies a threshold: beyond 272,000 input tokens, the request is billed twice as much on input and 1.5 times as much on output. DeepSeek’s price list has nothing of the kind.

For example, a request of 300,000 input tokens and 5,000 output. Without the threshold, Luna would bill $0.066; with it crossed, it costs $0.129. At DeepSeek: $0.096 at peak, $0.048 off-peak.

On very long requests, DeepSeek is therefore cheaper at any hour. It also writes longer in one go: 384,000 maximum output tokens against 128,000. Two caveats, though: its card shows a regression against V4 Pro on LongBench-V2, and Artificial Analysis measures both contenders at a perfect tie, 84%, on its long-context reasoning evaluation.


For consumers: the DeepSeek app against free ChatGPT

For most readers the question is not the API but “DeepSeek or ChatGPT?”. And the two models in this comparison are precisely the ones behind the two free offerings.

Free ChatGPT runs on GPT-5.6 Luna since the week of 6 August 2026. OpenAI added unlimited text conversations and a Think button for harder questions, “subject to anti-abuse guardrails”. Limits remain on files, images and other tools. The announcement does not say which reasoning level corresponds to the default mode, nor to the Think button — not a small detail, given the gap between settings measured above. We set out exactly what that tier contains in our comparison of Muse Glimmer against free ChatGPT.

DeepSeek V4.1 Flash is served on DeepSeek’s website and app since its launch on 10 September.

The first difference is visible on the home page. As of 16 September 2026, chatgpt.com lets you ask a question with no account, warning that conversations may be reviewed and used to improve the models. chat.deepseek.com opens on a login screen: phone number or email address, Google account or Apple account. For anyone searching for “DeepSeek free without sign-up”, the official site does not offer one.

The second difference is the only preference data that exists on these models, and it comes from French-speaking voters — a limit worth stating. On compar:IA, where 259,000 blind votes separated 83 models on 16 September, GPT-5.6 Luna scored 1,099 points and DeepSeek V4 Flash, the previous version, 1,063. The uncertainty ranges partly overlap — 1,070 to 1,111 on one side, 1,039 to 1,083 on the other — but compar:IA placed the two models in separate classes, the third and the fourth.

Read again on 26 September 2026, those numbers had moved — compar:IA runs on a continuous stream of votes. Across 260,094 votes and 77 models instead of 83, GPT-5.6 Luna rises to 1,101 points, V4 Pro to 1,047 and V4 Flash falls to 1,046: the two DeepSeek models are now level, the gap to Luna has widened, and V4.1 Flash still does not appear. On this ranking as on DeepSWE, DeepSeek’s new model remains unmeasured by third parties.

One reading caution: compar:IA itself warns that its ranking reflects subjective preferences, and measures neither the factual accuracy of answers nor overall model performance.


Your data: China, the United States, and the open-weights fire exit

On data, the two free offerings share something people forget: both reserve the right to use your conversations to improve their models.

At DeepSeek, the privacy policy updated on 10 February 2026 is explicit: personal data is collected, processed and stored in the People’s Republic of China, and used among other things to train the models, with an option to object. A clause dedicated to the European Economic Area, Switzerland and the United Kingdom appoints a representative, the Prighter group, and entities of the DeepSeek group handle storage, security and research. At OpenAI, the banner on account-free ChatGPT sums the rule up in one sentence: conversations may be used to improve the models.

But DeepSeek holds a card a closed model cannot play: its weights are public, under the MIT licence. Any organisation can download them, install them on its own servers, or buy them from the host of its choice — the seven providers listed by Artificial Analysis, with throughput from 106 to 539 tokens per second. In that case, it is the host’s data policy that applies, not DeepSeek’s.

Hence the paradox of this matchup: the Chinese model is the one whose place of processing you can most freely choose — provided you do not go through its vendor. We applied the same reasoning to European sovereignty in our comparison of Perplexity against Mistral’s Le Chat, and to safety and regulation in ChatGPT against Grok.

Self-hosting remains for well-equipped organisations. Even though only 8 to 16 billion parameters work on each token, all 552 billion have to be stored and loaded. For large-scale deployments, DeepSeek invites organisations with 2,000 GPUs and a storage cluster to get in touch, according to VentureBeat.


And in languages other than English?

Neither model card published by DeepSeek or OpenAI contains a non-English evaluation. The ten evaluations in the independent index are overwhelmingly English.

Two clues only. On its card, DeepSeek publishes MGSM, a multilingual maths evaluation: V4.1 Flash goes backwards, to 80.2 against 85.7 for V4 Flash. And on its own card, MultiLoKo, which tests local knowledge across languages, drops to 45.5 against 50.9 for V4 Pro’s base model. Add compar:IA, the only ranking fed by non-English votes, which puts Luna ahead of DeepSeek’s previous version.

None of that measures writing quality itself — register, idiom, accuracy of cultural references. For a writer, a lawyer or a customer-service team working outside English, that is the criterion that should decide. We flag the gap rather than fill it.


Which to choose

Take DeepSeek V4.1 Flash if you run agents that act — terminal, automation, chained tools —, if your calls come from an American working day, if your batch jobs can run at night or at the weekend, if your requests regularly exceed 272,000 tokens, if response speed matters, or if you want the option to host the model yourself or with a third party of your choosing. Our comparison of a free agent against a paid one applies the same reasoning to autonomous agents.

Take GPT-5.6 Luna if you process PDF documents, if your calls fall mostly in a European weekday morning, if your output volume is high and every written token counts, or if you want free access with no sign-up for everyday use. And set its reasoning level: at the default, it solves one coding task in nine.

For general consumer use, free ChatGPT is still the simplest: usable with no account, unlimited text once you sign up, and preferred by compar:IA voters over DeepSeek’s previous version. The DeepSeek app offers a newer, faster model, at the cost of a mandatory account and your exchanges stored in China.

On coding, do not count on a verification any time soon. We wrote on 17 September that it was worth waiting a few days: the official DeepSWE leaderboard was refreshed on 22 September and did not add V4.1 Flash. If its 74.2% is one day confirmed, the model will join the leading pack on code at a token price several times below Claude Opus 5 or GPT-5.6 Sol. As of 26 September, it remains a promise — credible given its two previous announcements, but a promise. If you have to decide now, lean on the one independent measurement that exists, Terminal-Bench 4.0, not on the vendor’s card.

You can also line these two models up against any others in our free comparator, which recalculates on every ranking update.


What to remember

  • DeepSeek edges ahead, narrowly. 40 against 38 on the Artificial Analysis index v4.3; five evaluations to DeepSeek, three to Luna, two ties.
  • The advertised rate is the off-peak rate, and peak hours are billed in UTC: Monday to Friday, 01:00–04:00 and 06:00–10:00.
  • Geography decides. An American working day never touches those windows, so it is billed off-peak from end to end. A European working day loses two to three hours a day to them.
  • DeepSeek writes 2.2 times more tokens. Sitting the same exam costs it $477 at peak rates against $320 for Luna; about $238 at off-peak rates.
  • On an American working day, DeepSeek stays cheaper until it writes 2.28 times more than Luna; on a European one the threshold falls to 1.77 or 1.58. The measured ratio is 2.17 — which puts the same model on either side of the line depending on the continent.
  • The 74.2% on coding is still unverified by the official DeepSWE leaderboard, refreshed on 22 September 2026 without V4.1 Flash — though DeepSeek’s two previous announcements were confirmed there to within a point.
  • At the API’s default setting, Luna solves 11.3% of coding tasks, against 67.2% at maximum.
  • DeepSeek answers ten times faster than Luna at maximum: 12.6 seconds against 131, reasoning included.
  • Since 14 September, V4 Pro is no longer served: its requests go to V4.1 Flash, cheaper, stronger at agentic work, but less knowledgeable.
  • Free ChatGPT works without an account; DeepSeek requires sign-up and stores your data in China. Both may use your conversations to improve their models.
  • Open weights are DeepSeek’s real fire exit: seven providers, up to a 4.9× price spread, and the option to keep everything in house.