On 12 August 2026, xAI released Grok 4.6. The next day, 13 August, Google released Gemini 3.7 Flash. Twenty-four hours between two models aimed at the same work: code, long-running agents, chained tasks.

On the software engineering evaluation both vendors put forward, DeepSWE v1.1, the gap is six tenths of a point. On the bill for everyday use, it is more than double. And on a single long-context request, it climbs to five times.

That is the gap this comparison set out to measure, because it is the one that actually decides.


Results: the verdict in brief

Gemini 3.7 Flash takes the lead on available context, accepted formats, price, and four of the five categories in our ranking. Grok 4.6 keeps the edge on pure code, on the Artificial Analysis composite intelligence index, and on one thing Google lost along the way: nothing in its documentation asks a European administrator to accept a routing warning.

But the essential point is in neither announcement. Gemini 3.7 Flash hallucinates more than the version it replaces: 64.5% against 55.6%. Grok admits it does not know only two times in three. On image reading, the gap reaches 69% against 20% on object detection. And on knowledge work it is Grok that pulls away, with a 228-point Elo lead — even though it was announced as a coding model.

As for speed, which every comparison turns into an argument: the published figures vary by a factor of fifty from one source to another. We explain below why, and why none of them transfers to your setup.

Neither generates images or sound. Neither has been evaluated by its vendor in any language other than English. And both pose, each in its own way, a data problem for a European user.

Here is the detail.


Where these numbers come from

Method before figures rather than in a footnote, because a comparison is worth only what its sourcing is worth. Every figure in this article can be replayed line by line:

  • All cost calculations were redone from the rates published by both vendors, with the assumptions detailed below so you can check them yourself.
  • Performance cross-checks four independent series — Artificial Analysis, DeepSWE v1.1, SWE-bench Vals and our own ranking — never a single leaderboard. When they diverge, we say so instead of picking the one that suits the argument.
  • Production figures — latency, throughput, error rate, cache hit rate, effective cost — are production logs, attributed in the body of the article to whoever published them.
  • Ranks by category and by reasoning level come from our public ranking, read on 30 August 2026.
  • Incapabilities and caps are quoted from the official documentation, including the warnings the vendors write about their own products.

Two points are not measurable as things stand, and we flag them rather than fill them: quality outside English, which neither vendor evaluates publicly and for which no quantified comparison exists, and Grok 4.6’s rate limits, which xAI has not published.


The two spec sheets, side by side

Grok 4.6 Gemini 3.7 Flash
Vendor and date xAI — 12 August 2026 Google — 13 August 2026
Context 500,000 tokens 1,048,576 tokens (2.1 times more)
Maximum output not published 65,536 tokens
Accepted inputs text, image text, image, audio, video, PDF
Output text only text only
Knowledge cutoff 1 February 2026 not published
Tools function calling, structured outputs function calling, Google Search, computer use (preview)
Price input / output $2.00 / $6.00 per million $0.75 / $3.75 until 31 Dec 2026
Cached input $0.50 per million —
Beyond 200,000 tokens $4.00 / $12.00 on the entire request rate unchanged

Two remarks the spec sheet does not make out loud. First: the word “multimodal” here describes what goes in, never what comes out. Neither produces an image or a sound. Second: Gemini swallows PDFs, audio and video where Grok is limited to text and images. If your work consists of having a model read documents, the question is settled before performance even comes up.


What 500,000 and 1,048,576 tokens actually represent

The context figure is the most quoted and least understood number on either spec sheet. Here is how to translate it.

A token is not a word. In English the conversion is favourable — roughly three words for every four tokens — so a million-token window holds around 750,000 words: several thousand pages, or a mid-sized codebase in full. In French, where accents, elisions and punctuation cost extra, the same window holds between 500,000 and a million words. Worth knowing if your corpus is not in English: the same context window is effectively smaller.

What that changes in practice: Grok’s 500,000 tokens are amply enough for a file, a case folder, a long exchange. Gemini’s million becomes useful in one scenario, but an increasingly common one — the agent that has to keep a whole project in view from one end of a long task to the other, without you having to cut and re-stitch the pieces. If your work fits in a file, the context gap does not concern you. If it means making a model reason over an entire repository, it decides everything.

One useful clarification: the context window is not memory. It is emptied on every request. A model with a million tokens of context does not remember you better from one session to the next — it simply reads more things at once.


Performance: two independent measurements, two opposite verdicts

This is the most interesting result of the comparison, and it comes down to one sentence: depending which source you consult, a different model wins.

Measurement Grok 4.6 Gemini 3.7 Flash
Intelligence index (Artificial Analysis) 61 56 (high reasoning)
DeepSWE v1.1 65.9% 65.3%
Terminal-Bench v3.0 26% (GPT-5.6 Sol: 34.6%) not published
FrontierCode 61.3% 43.6% (against 34.4 on 3.6)
SWE-bench (Vals) 95.60% not published
GDPval-AA v2 (Elo) — knowledge work 1753 1525 (228 points apart)
Terminal-Bench 2.1 88.4% 85.8%
Long context (GDM-MRCR v2 at 128K) not published 97.0%
AutomationBench not published 30.4% (against 17.0 on 3.6)
WebDev Arena (Elo) not published 1588 (against 1538 on 3.6)

On the composite index, Grok leads by five points. On the software engineering evaluation both publish, it leads by six tenths — in other words, nothing.

One line stands out sharply, though: GDPval-AA v2, which measures knowledge work rather than code, gives 1753 to Grok against 1525 to Gemini. Two hundred and twenty-eight Elo points is no longer a measurement gap, it is a difference in kind. That result matches a remark several practitioners made: Grok 4.6 was announced as a coding model, when its real strength lies elsewhere — analysis, synthesis, reasoning over documents. If your need is to make a model think about documents rather than write functions, the ranking inverts.

At the other end, Gemini is the only one to publish a long-context retention measurement: 97.0% on GDM-MRCR v2 at 128,000 tokens. xAI publishes none, which leaves the most legitimate question about a 500,000-token window unanswered — does it hold all the way?

The table stays full of holes: each vendor picks the evaluations that flatter it and stays quiet about the rest. That is the structural limit of any model comparison, and it is why a fifth source is needed.


Vision: the widest gap in the comparison

This is the dimension both announcements barely mention and comparisons forget, even though it decides entire use cases — reading a screenshot, extracting a table from a scanned document, controlling an interface.

Roboflow’s vision evaluations give an unambiguous result:

Vision evaluation Grok 4.6 Gemini 3.7 Flash
Overall average 68% 85%
Object detection 20% 69%
Image reasoning 62% 83%

Seventeen points of average, and above all a ratio of more than three to one on object detection: 20% against 69%. At that level it is no longer a matter of preference. A workflow that rests on reading images — inventory, quality control, document extraction, interface automation — is not feasible with Grok 4.6 today.

Remember that Grok accepts images as input: it sees them. It simply interprets them far less well. And since it reads neither PDF, nor audio, nor video, the gap on anything that is not plain text widens further.


The field figures: latency, throughput, error rate

Composite indices say nothing about what using them feels like. These measurements come from production logs and service data, not press releases — and this is where the two models stop resembling each other.

Measured in real use Grok 4.6 Gemini 3.7 Flash
Throughput 94 tokens / second 93 tokens / second (median)
Average latency 0.62 seconds 1.85 seconds (nearly 3 times more)
Errors on tool calls 0.08% not reported
Errors on structured outputs 4.00% — but 0.00% on the zero-retention endpoint not reported
Published rate limits 150 requests/s, 50M tokens/min not published for this model

A necessary warning before going further: on these two models, speed is the least reliable number in circulation. The logs in the table above give near-identical throughput, around 93 to 94 tokens per second. Other comparators publish 340 tokens per second for Gemini in high reasoning against 85.8 for Grok, a ratio of nearly four to one in the other direction. On time to first token, the spread between sources reaches a factor of fifty: 0.62 seconds here, 32.3 seconds there, for the same Grok 4.6.

Those figures are not wrong, they do not measure the same thing. A reasoning model goes through a thinking phase before emitting its first visible token: depending on whether you time the start or the end of that phase, you get half a second or half a minute. Add the requested reasoning level, the provider you go through, and the time of day. The practical conclusion is that no published speed figure on these two models transfers to your setup, including these. What to take away is not a value but a dominant variable: the reasoning level decides everything else.

That said, one thing is stable from source to source: the ratio between time spent and result obtained. On the Artificial Analysis index tasks, Gemini finishes in about 1.7 minutes per task. One developer published a log that sums up the experience well: on the same task, DeepSeek V4 Pro took 12 minutes 02 for $0.12 with a bug, while Grok 4.6 took 3 minutes 18 for $1.41 with none. More expensive, much faster, clean result — exactly the trade Grok offers.

On Gemini’s side, latency depends on where you go through: Google AI Studio serves 124 tokens per second at 1.74 seconds, while Vertex drops to 84 tokens per second at 2.02 seconds. Same model, two front doors, different performance — an operational detail product pages never mention.

Time to first token, meanwhile, depends on the reasoning level, and the amplitude is considerable: 0.74 seconds at low effort, 9.83 seconds at high effort, and up to 49.5 seconds at the 95th percentile. For an interface where someone is waiting for an answer, that is not a nuance: it is the difference between a usable tool and an abandoned one.

One detail we have seen nowhere else, and worth the detour: at Grok, the error rate on structured outputs falls from 4.00% to 0.00% on the zero-data-retention endpoint. In other words, the entry point that best protects your data is also the one that fails least. We do not explain why — but if you generate JSON at scale, that information is worth money.


The subject neither vendor talks about: hallucination

Neither launch announcement contains a line about hallucination. The independent measurements do exist, and they are harsh on both.

Gemini 3.7 Flash hallucinates more than the version it replaces. Artificial Analysis measures a rate of 64.5%, against 55.6% for Gemini 3.6 Flash. Read that sentence twice: accuracy went up, and hallucination went up at the same time. The explanation lies in the model’s tuning, which favours the confident answer over the cautious one. An update that makes a model both more accurate and more inventive is ambiguous news, and nobody wrote it down.

Grok 4.6 admits it does not know two times in three. Its non-hallucination rate is measured at 65.7%: when it does not know, it says so about two times in three, and invents the third. That is precisely the evaluation line missing from xAI’s launch post.

What that means concretely was illustrated by a production case reported publicly: a European car dealership’s customer service assistant confidently confirmed support for brands absent from its database, because the help centre stated somewhere that “we support all models”. The model had not lied in the strict sense — it had filled a gap with confidence.

There is a second figure to set beside it. On τ³-Banking, the evaluation closest to real support work, Grok 4.6 ranks in the top two with 50.7%. That is the best on the board, and it still fails one conversation in two. The practitioners’ conclusion is unanimous and worth repeating: for customer service, the choice of model is the least decisive decision. What decides is the retrieval boundaries, the escalation rules and the confidence thresholds — and neither training provides them.


What our own ranking says

Our AI ranking aggregates around twenty public sources and scores models by use category. Read on 30 August 2026, at high reasoning level for both:

Category Grok 4.6 Gemini 3.7 Flash
General-purpose 107th 83rd
Conversation 198th 135th
Code 50th 60th
Cost-performance 152nd 89th
Cloud agentic 59th 47th

Four categories out of five to Gemini, one to Grok — and it is code. That result contradicts the Artificial Analysis index, which gives Grok a five-point win. Both measurements are valid: the composite index weights pure reasoning, our aggregation weights the average of around twenty different protocols. When two honest methods diverge, it means the real gap is thin and depends on your use. The detail is visible on the code category and on cost-performance.

One caution about these ranks: several hundred models are ranked in each category, and a single family appears there under several variants. A 50th place is not a bad mark, it is a position in a very dense population.


The setting that costs more and ranks worse

Here is the finding we have seen nowhere else, and it has a direct consequence on your bill.

Grok 4.6 offers four reasoning levels: low, medium, high and xhigh. Intuition says the harder you push, the better it gets. Our readings say the opposite, in all four categories where both levels coexist:

Category Grok 4.6 "high" Grok 4.6 "xhigh"
General-purpose 107th 162nd
Conversation 198th 265th
Code 50th 94th
Cost-performance 152nd 260th

Four times out of four, the harder setting ranks worse. And since reasoning tokens are billed as output, xhigh is also the more expensive setting. We do not claim to explain why — an aggregation of protocols is not a diagnosis. But the observation is enough to give practical advice: if you use Grok 4.6, test high before paying for xhigh, and measure on your own workload. Prudence says treat that level as an option to validate, not as a superior default. We found the same counter-intuitive pattern on GPT-6 Astra against Claude Fable 5.1, where the maximum setting is the worst buy in either catalogue.


Price: the calculation nobody does

Comparisons generally settle for lining up the per-million-token rates. That is not what you pay. Here are three complete scenarios, with the assumptions written out so you can replay them.

Scenario A — everyday coding assistant. 50 requests a day, 30,000 input tokens and 3,000 output per request, 30 days.

Scenario Grok 4.6 Gemini 3.7 Flash Gap
A — everyday use, current rates $117.00 / month $50.63 / month 2.3 times
B — same use, after 1 January 2027 $117.00 / month $101.25 / month 1.2 times
Measured cost per task (AA index) $0.84 $0.40 (high) — $0.26 (medium)
C — one 250,000-token request $1.06 $0.21 5.1 times

Scenario A gives the everyday gap: more than double. The measured cost-per-task line confirms it from another angle — $0.84 against $0.40, a two-to-one ratio, rising to more than three if you drop Gemini one reasoning notch with no noticeable quality loss. That is in fact the best-value setting in the comparison: $0.26 per task at medium reasoning.

Scenario C deserves its own section, because it is not explained by the difference in advertised rates.


The 200,000-token threshold: Grok’s billing trap

At xAI, the grid changes tier at 200,000 tokens. That is not surprising in itself — many vendors bill long context more. What is more surprising is how: beyond the threshold, it is not the overage that switches to the long-context rate, it is the entire request.

Take scenario C again. A request of 250,000 input tokens and 5,000 output:

  • If the threshold did not exist: 250,000 × $2/M + 5,000 × $6/M = $0.53
  • With the switch: 250,000 × $4/M + 5,000 × $12/M = $1.06
  • The same request at Gemini: 250,000 × $0.75/M + 5,000 × $3.75/M = $0.21

Crossing the threshold doubles the bill, and takes the gap with Google to 5.1 times. For an agent that re-reads an entire codebase or a long document set, the difference is no longer theoretical: it shows up at the end of the month.

First practical consequence: watch your request sizes if you use Grok. Going from 199,000 to 201,000 tokens does not cost 1% more, it costs 100% more. OpenAI applies the same mechanism at 272,000 tokens on GPT-6 Astra, as we detailed in that comparison — the trap is becoming a pattern, not an xAI quirk.


Caching: the lever that almost cancels the price gap

It would be dishonest to stop at scenario A without saying what can reverse it, and this is the second calculation we have seen nowhere.

Grok bills cached input at $0.50 per million, a quarter of the normal input rate. For an agent that resends the same context every turn — the system instructions, the reference documentation, the file it is working on — most of the input is repeated, and therefore cacheable.

Take scenario A again, this time with input fully cached:

  • Input: 50 × 30,000 tokens at $0.50/M = $0.75 a day (against $3.00 without cache)
  • Output: 50 × 3,000 tokens at $6/M = $0.90 a day, unchanged
  • Total: $1.65 a day, so $49.50 a month

Against Gemini 3.7 Flash’s $50.63, the 2.3-times gap vanishes outright. Grok even edges very slightly ahead.

That theoretical calculation is confirmed in the field, which is rare: on real workloads, a cache hit rate of 90.3% has been recorded, for an effective cost of $0.72 per million input tokens — nearly three times less than the $2 advertised rate. Our lower bound was therefore not a textbook case.

Two honest caveats all the same. First: only the part repeated identically from one request to the next is eligible — 90% is an excellent result, not a guaranteed floor, and it assumes an architecture built for it. Second: Google does not publish an equally legible cache rate for this model, which does not mean no optimisation exists on its side, but that it is not comparable as things stand.

Worth noting in passing: xAI raised the cache rate from $0.30 to $0.50 between version 4.5 and 4.6, a 67% increase on the main saving lever. The advertised price did not move; that one did.

What this calculation does establish is solid: the price gap between the two models depends more on your architecture than on their price lists. A well-designed agent that reuses its context almost erases Google’s advantage. A single bulky call, conversely, takes it to 5.1 times. Same pair of models, two unrelated invoices.


Gemini’s advantage has an expiry date

The $0.75 and $3.75 rate is not Gemini 3.7 Flash’s rate. It is an introductory rate, valid until 31 December 2026. On 1 January 2027 it goes to $1.50 and $7.50 — exactly double.

What that does to the comparison is spectacular, and it is scenario B in the table: the same everyday use goes from $50.63 to $101.25 a month. The gap with Grok, which stays at $117, falls from 2.3 times to 1.2 times.

In other words: a reader choosing Gemini today for its price should know they are comparing a promotional rate with a permanent one, and that the decision replays in four months. That is the kind of detail that never appears in a price table, and yet decides an annual commitment. The sequence is not isolated: it belongs to the price war we traced in DeepSeek V4.1 Flash against GPT-5.6 Luna.


What they cannot do

A useful comparison also says where both contenders stop. On this ground, the official documentation is franker than the press releases.

Neither generates images or sound. Google’s card is explicit: audio generation, image generation and the Live API are not supported by Gemini 3.7 Flash. Grok 4.6 accepts text and images as input, and returns text only. The word “multimodal” qualifies the input, not the output.

Grok reads neither PDF, nor audio, nor video. That is probably the most concrete difference in the comparison for ordinary professional use: summarising a PDF report or a recorded meeting is a mundane use case, and Grok 4.6 does not cover it natively.

Gemini’s computer use is in preview, and Google attaches a warning worth quoting as written: suggested actions may be inappropriate or dangerous, adversarial content can make them malicious, and the vendor recommends supervising closely and not using it for tasks involving critical decisions, sensitive data, or actions whose serious errors cannot be corrected. A vendor who writes that about its own product deserves to be read literally.

You cannot combine Google Search and function calling in the same request. The API documentation states it: search tools do not mix with non-search tools in the same call; multiple tools are admitted only if they are all search tools. You therefore have to split into two calls, at a cost in latency and tokens. Grounding with search is also capped at one million requests a day.

Moving to Gemini 3.7 Flash requires touching your code. The temperature, top_p and top_k parameters have been removed: an existing integration that uses them has to be restructured, this is not a simple model identifier change. Several developers also reported API key creation being refused on fresh accounts, with an unhelpful message — “the request is suspicious, please try again”.

And the documentation lagged behind the product. On 14 August, the day after release, the Gemini API’s public model index, the changelog, and the pricing and limits pages still featured Gemini 3.6 Flash. An available model whose documentation has not caught up means time lost on integration — the kind of friction that appears in no comparison table.

Reasoning tokens inflate the output bill more than you think. In the example trace Google provides, one response counted 297 reasoning tokens for 171 visible output tokens: nearly two thirds of the “output” line on your invoice corresponds to reasoning you will never read. Our cost calculations, which count only visible output, are therefore low estimates as soon as reasoning is on.

xAI has not published Grok 4.6’s rate limits. For reference, version 4.5 was given as 150 requests per second and 50 million tokens per minute on production tiers. Designing a production workload on an unpublished figure is a gamble, and you should know it before committing.


Your data: two steps backwards, of different kinds

This is where this comparison diverges most from others, because the question does not arise the same way depending on where your data has to live.

At xAI, training on your conversations is on by default. The published policy provides that your prompts, the content you upload and the model’s responses may be used to train and fine-tune Grok. It is not hidden, but it is a setting to switch off, not to switch on. If you delete a conversation or an account, the data is queued and generally removed within thirty days, barring legal or security obligations.

At Google, the new release is a compliance regression. Gemini 3.7 Flash is available only in the so-called global region, with no data residency. An administrator selecting it from a European deployment has to accept a warning saying that traffic goes to the global endpoint. Yet the previous version, Gemini 3.5 Flash, remains available in European regions — Belgium, the Netherlands, Finland, Poland.

The conclusion is unusual and needs writing as it stands: for an organisation under a localisation constraint, the update is a step backwards, and the right decision is to stay on the earlier model. This concerns any company with European operations, wherever its head office sits. We examined the same tension between capability and sovereignty in our comparison of Perplexity against Mistral’s Le Chat, and the regulatory calendar in ChatGPT against Grok.


Languages other than English: what nobody measures, us included

Support for languages other than English is the criterion comparisons skip most systematically. Here we have to flag a hole, and it is not of our making.

Neither xAI nor Google has published a non-English evaluation for these two models. The evaluations put forward — DeepSWE, SWE-bench, GDPval, AutomationBench, WebDev Arena — are overwhelmingly English-language, and the independent leaderboards, ours included, aggregate those same evaluations. No public comparative measurement therefore exists to date.

Extended use obviously produces impressions, one way or another. But an impression is not a measurement, and this comparison refuses to convert the second into the first: ranking two models on the quality of their French, German or Japanese today would be dressing a preference up as a result. The hole is itself information: two frontier models released a day apart in August 2026, and not one public figure on how they behave outside English.


Where to actually use them, and at what price

A confusion comes up often: the per-million-token rates only concern the API. For personal use, both are sold differently.

Grok 4.6 on the consumer side is used on grok.com, accessible with a Google account or just an email address — an X account is not required. Access is sold as a monthly subscription with usage caps: around $10 for the light plan, $30 for the middle one, $300 for the highest. On the developer side, the model is exposed in the xAI API, Grok Build and Cursor, and reachable through OpenRouter, Vercel and Cloudflare.

Gemini 3.7 Flash is used through the Gemini API, Vertex AI and Google Cloud’s agentic platforms, and appears in Google’s consumer plans.

The rule that follows is simple: if you are a regular but non-developer user, the subscription is almost always cheaper than the API, because it pools usage peaks. The calculations in this article apply to an application integration, not to individual use in a chat interface.


Which to choose for your situation

You write code and the budget follows. Grok 4.6. It is the only category where it beats Gemini in our ranking, it leads the Artificial Analysis index by five points, and its SWE-bench result measured by Vals is the highest it has published. Watch your request sizes and exploit the $0.50 cache.

You have documents analysed, not code written. Grok 4.6, and more clearly still: 228 Elo points ahead on GDPval-AA v2, the evaluation that measures knowledge work. That is the paradox of this model, announced for code and better elsewhere.

Your work goes through images. Gemini 3.7 Flash, without hesitation. 69% against 20% on object detection: that is not a comfort gap, it is the difference between a workflow that works and one that does not.

You are building an agent that runs long, over many documents. Gemini 3.7 Flash. Twice the context, native reading of PDF, audio and video, a better place in agentic work and cost-performance with us, and an advertised rate half as high — until the end of the year. Caveat: if your agent reuses the same context every turn, Grok’s cache erases that advantage and the decision replays.

You are under a data localisation constraint. Neither, as things stand. Grok trains by default on consumer conversations; Gemini 3.7 Flash offers no European residency. The compliant answer today is called Gemini 3.5 Flash, in a European region.

You want a model to read a PDF or a recorded meeting. Gemini, no discussion. Grok does not do it.

You are still hesitating and you do not write code. Take a subscription rather than an API key, test both for a month, and keep in mind that these two models are twenty-four hours and six tenths of a point apart on the common evaluation. At that level of proximity, your use case decides better than any leaderboard — ours included. You can line them up yourself in our free comparator, which recalculates on every ranking update.


What to remember

Two models released a day apart, six tenths of a point on the common evaluation, and a factor of five on a single long-context request. Most of the decision does not turn on performance, but on three lines the product pages do not highlight: a billing threshold at 200,000 tokens at Grok, an expiry date of 31 December 2026 at Gemini, and a European data residency that Google’s latest version lost along the way.

And there is one result we did not expect, which deserves checking against your own usage: at Grok 4.6, the hardest reasoning setting is also the most expensive and the worst ranked, in all four categories where we can compare them.