On 27 August 2026, Google shipped Gemini Omni 1.1 Flash. A few weeks earlier, on 6 August, Alibaba had opened Wan 3.0 in public beta. Two video generation models, two giants, and in our ranking thirty-five hundredths of a point between them — 99.46 against 99.11.

Which is to say, nothing.

Except that neither announcement says the same thing as the documentation behind it. Google sells forty seconds: they are three to ten second pieces stacked end to end. Google sells 4K: it is an upscale. And the feature that makes the model interesting — editing video you shot yourself — is switched off in the United Kingdom and across the European Economic Area.


The verdict in brief

On the length of a single shot: Wan 3.0, no argument. Thirty seconds in one pass, with sound, against three to ten seconds at Google that then have to be joined.

On the edit loop: Gemini Omni. It keeps the state of the video from one turn to the next and is driven in plain language. It is an iteration tool, where Wan is a production tool.

On the price of a delivered minute: Google, by about 31 %. But not for the advertised reason — the saving does not come from the output price, it comes from the draft tier at $0.03 per second.

On what you can feed it: Wan 3.0, by a wide margin. A PDF, a deck, a spreadsheet, a web page. Google caps out at three clips of three seconds and refuses YouTube addresses.

On availability: Google loses its flagship feature in the United Kingdom and the European Economic Area — and keeps it in the United States. Editing video you shot yourself is unavailable in the EEA, in Switzerland and in the United Kingdom. This is the one criterion where your country changes the answer, and we come back to it in detail.

On on-screen text in English: Google, and this reverses the French edition of this comparison. English is the only language Google says it has evaluated. Wan announces twelve, and admits in the same breath that its text accuracy is not where it wants it.

And on the honesty of the announcements: Alibaba, to everyone’s surprise. The Chinese vendor writes in its own press release that its audio and its burnt-in text are not at the level it wants.


Where these numbers come from

Three families of data, kept apart from end to end.

Vendor announcements and documentation supply capabilities, caps, prices and restrictions. They are useful for what a product does and refuses to do — and it is precisely by setting one against the other that most of this article appeared: what a press release claims is not always what an API page allows.

Practitioners’ reports come from people who generated sequences with each of the two models and published what they got: whether a character holds up over time, how materials behave, dialogue quality, failures. They are attributed by name in the text.

Ranks come from the Le Recul AI ranking snapshot of 9 September 2026, which aggregates around twenty public sources.

One reading principle, applied throughout: an observation is not a measurement. When image quality is judged, it is attributed to whoever judged it. When a figure is calculated, its assumptions are written out so you can redo it.

One correction, made during preparation and flagged here because it nearly slipped through: a secondary source claimed Alibaba had put Wan 3.0 online with no announcement at all. That is false. The vendor published on its official blog on 6, 7 and 13 August 2026. Those posts are what count in this article.


The two contenders

  Gemini Omni 1.1 Flash Wan 3.0
VendorGoogle DeepMindAlibaba Cloud, Tongyi Lab
Date27 August 2026public beta, 6 August 2026
Duration in one pass3 to 10 s4 to 30 s
Maximum cumulative duration40 s by stacking30 s, plus extension
Resolutions360p, 720p, 1080p, 4K — the two highest upscaled480p, 720p, native 1080p, no 4K
Frame rate24 frames/snot published, no 60 frames/s
Soundgenerated, voice editing not supportedgenerated with the image in one pass
Inputstext, image, audio, video (3 clips of 3 s max)text, image, audio, video, documents, web pages
Maximum references3 clips20 items
On-screen textEnglish fully supported, other languages not evaluated12 languages announced
Editing an uploaded videoblocked in the EEA, Switzerland and the United Kingdomno equivalent restriction documented
WatermarkSynthID on every outputnot documented
Open weightsnono — last open weights: Wan 2.2

”Forty seconds”: what Google sells and what the API delivers

The announcement is spectacular. The documentation much less so.

Every call to Gemini Omni 1.1 Flash generates between three and ten seconds. The forty seconds are a cumulative total: you append slices to the end of a clip you already produced, until you hit the ceiling. XenoSpectrum says it without padding — the idea of a forty-second video generated in one go is not accurate.

Three constraints come with it, and they weigh on real work:

  • Extension only works at the end of a clip. No extending at the start, no inserting in the middle. If the missing shot is the middle one, you start again.
  • Each pass is billed separately, at whichever tier you pick. There is no extension surcharge, but no discount either.
  • The model only analyses the last ten seconds to continue — already a clear improvement, since the previous version looked at one, but still a short memory for a forty-second sequence.

The progress is real. The wording invites you to believe something else.


”4K”: the word that does not mean what you think

Same mechanism, one notch more awkward.

Gemini Omni’s 4K is not generated, it is upscaled. And the spec sheet goes further than the announcement lets on: of the four tiers on offer, the two highest are upscales. So 1080p itself is not native.

The workflow Google recommends is consistent with that architecture: explore in 360p, validate in 720p, export in 1080p or 4K. It is an honest production pipeline — as long as you know the export adds no detail, it interpolates it.

Wan 3.0 does the opposite and owns it: native 1080p as its ceiling, no 4K tier at all. You get fewer pixels, but the ones you get were generated.

For a social media post, the gap is theoretical. For a client delivery where someone is going to zoom in, it stops being theoretical.


Thirty seconds in one take, and it is true

This is Alibaba’s central argument, and it holds.

The official announcement is clear: up to thirty seconds in a single pass, with smart duration recommendation, and sound generated at the same time as the image. A clip comes back with audio, not silent. The hands-on review by Wan30 confirms the real range — four to thirty seconds per generation.

The practical consequence matters more than the duration itself. Thirty seconds in one pass is a continuous take. Three ten-second shots stacked together is an edit — with three chances for a character to change jacket, for the light to shift, for a camera move to contradict itself.

And on this precise point nobody can settle it with figures: Google publishes no consistency measurement and no failure rate for scene extension. It is the most visible hole in the announcement.


What Wan accepts as input, and nobody else does

Here is the feature people talk about least and that changes the most for professional work.

On top of text, images, audio and video, Wan 3.0 accepts documents and web pages as input. The official list is specific — doc, xls, ppt, pdf, txt, key, pages, numbers, md — within a limit of one file or one link, one hundred megabytes, fifty pages. Total reference items can go up to twenty.

In other words: a sales deck, a product sheet in PDF, a page from a website can become a video without the intermediate step of “translate all this into visual instructions”.

Facing that, Gemini Omni caps video references at three clips of three seconds, ignores the sound of those references, refuses YouTube addresses, and cannot reason from one reference clip to another.

What you can feed in Gemini Omni 1.1 Flash Wan 3.0
Textyesyes
Imageyesyes
Reference video clips3 maximum, 3 s eachincluded in the 20 references
Uploaded video to edit10 s maximum — blocked in the EEA and the UKyes
Sound of those referencesignoredtaken into account
Documents (pdf, ppt, doc, xls, txt, md, key, pages, numbers)noyes — 1 file, 100 MB, 50 pages
Web page by its addressnoyes
YouTube addressrefusednot documented
Total number of references320

This is not a difference in power. It is a difference in trade: Google is building a director’s tool, Alibaba a marketer’s tool.


Price per second, tier by tier

Tier Gemini Omni 1.1 Flash Wan 3.0
Draft$0.03/s (360p)$0.05/s (480p)
720p$0.10/s$0.10/s
1080p$0.15/s (upscaled)$0.20/s (native)
4K$0.30/s (upscaled)not offered

Three observations the grid alone does not give you.

At 720p it is a dead heat — ten cents a second on both sides. The most used tier on the market settles nothing.

At 1080p, Google is 25 % cheaper, but its pixel is upscaled. So you are choosing between paying less for interpolated output and paying more for generated output. That is not the same thing as “cheaper”.

And this version is not a price cut. XenoSpectrum makes the point: at $0.10 per second in 720p and $0.30 in 4K, Gemini Omni 1.1 Flash matches Veo 3.1 Fast exactly. What changed is not the price of the deliverable, it is the price of the attempt. We watched the same mechanism at work when DeepSeek set off an API price war — the headline rate moves less than the shape of the bill.


The real calculation: what a delivered minute costs

This is the figure nobody publishes, because it forces you to admit how much you throw away.

Assumptions, written out so you can replay them: delivery in 1080p, and three drafts for every shot kept — a realistic rate in video generation, where the first take is rarely the right one.

One delivered minute in 1080p Gemini Omni 1.1 Flash Wan 3.0
Shot breakdown required6 shots of 10 s2 shots of 30 s
Drafts (3 per shot)6 × 3 × 10 s × $0.03 = $5.402 × 3 × 30 s × $0.05 = $9.00
Final renders6 × 10 s × $0.15 = $9.002 × 30 s × $0.20 = $12.00
Total$14.40$21.00
What you get6 generations to splice2 continuous takes

Google is about 31 % cheaper. And the lever is not the one you would guess: it is not the output price, it is the draft tier at three cents. Build Fast with AI puts it better than any price grid — the model is not cheaper in absolute terms, it makes experimenting cheaper.

But read the last row of the table before concluding. Google’s $14.40 buys six separate generations that will have to be joined, whose consistency no published figure guarantees. Alibaba’s $21 buys two continuous takes. The price is not measuring the same thing on each side — and six bad joins cost more than six dollars and sixty cents.


What people who have run Wan 3.0 report

The observations published by Video Generator cover five sequences: a ceramic mug product shot, a vertical headphone unboxing, a twelve-second market scene, a physics test on a paper lantern, and a twenty-second cinematic sequence in an alley in the rain.

In the model’s favour:

  • The fifteen-second mug shot is judged almost usable on the first attempt.
  • Consistency holds: “his jacket stayed the same colour, his face stayed recognisable”.
  • The physics convinces: “fabric folds like fabric. Water splashes and falls back.”
  • Ambient sound is judged good enough for a working edit.

Against it:

  • Hands remain a problem — the sector’s oldest defect is not solved.
  • Dense crowds soften at the edges of the frame.
  • Dialogue sounds slightly processed, which makes it hard to use as is.
  • Multi-shot sequences feel like a collage rather than a continuous take.

Compared with its rivals by the same tester: against Veo 3, Wan “trades blows rather than loses”, while staying less photorealistic; against Kling and Runway, it is “a clear step up on consistency and camera control”. Against Sora, they are two philosophies — spectacle on one side, repeatability and price on the other.

We had already measured that shift when we set Kling 3.0 against Veo 3.1: the playing field has moved from how beautiful a shot is to how reliably you can reproduce it.


What people who have run Gemini Omni report

The score is good — 8.5 out of 10 at Build Fast with AI. The reservations are specific, and one of them is severe.

Three recurring weaknesses: lip-sync drift, movement that floats, and trouble as soon as there are several speakers to synchronise.

But the most awkward finding concerns the flagship feature: scene consistency degrades across successive edits. In other words, the model is least reliable exactly where it is supposed to excel — iteration.

Google does not hide it in its technical documentation, where it states that full consistency across edits, complex motion and perfectly accurate text remain difficulties. The wording is honest. It simply is not in the press release.

One additional warning is worth knowing if you are planning a production chain: conversational memory promises no pixel-perfect preservation. A face, a garment, a framing can change without being asked. The documented workaround fits in one sentence to append to every instruction: “keep everything else identical”.


The silent failure: zero tokens, no message

Here is the most expensive trap in time, and it belongs to Gemini Omni alone.

When the model refuses a clip, it does not always say so. A video-to-video edit can finish instantly with zero tokens produced and an empty output, with no explicit error message. You think you have a bug in your code; it is a refusal.

Three causes explain almost all of these refusals:

  1. The regional lock on uploaded video — we come to it next.
  2. The model only reliably edits its own generations. One user quoted reports that every one of their clips was rejected, and that the only videos accepted were the ones the model had made itself.
  3. An anti-deepfake rule, applied globally, that refuses to make a photograph speak from an audio file.

Generic messages add to the confusion. One developer sums up the general feeling by writing, about a block in India, that it would have been better to show a real error message.

The workaround fits in one workflow rule: generate inside the model first, edit afterwards. Which, incidentally, turns a limitation into a business model.


The point that decides — and it changes with your country

We checked this on six independent sources before writing it, because it changes the verdict. And unlike the French edition of this comparison, it does not change it the same way for every English-speaking reader.

Editing and extending a video you uploaded are unavailable in the European Economic Area, in Switzerland and in the United Kingdom. The restriction sits in Google’s API documentation, so it is not just the consumer app: there is no programmatic route either. Personal avatar creation and uploading images containing minors fall under the same territorial blocks.

So the answer splits:

  • If you are in the United Kingdom, in Ireland or anywhere in the European Economic Area, this is decisive. Note that the United Kingdom appears by name: this is not an EU-only rule that stops at the Channel.
  • If you are in the United States, Google’s documented list does not cover you, and the model keeps the feature it was built around. User reports do add India and certain American states, including Texas and Illinois — but those are reports, not a published restriction, and we label them as such rather than promoting them to documentation.

What remains possible under the block: generate a video with the model, then extend and retouch it. What is blocked: feeding it footage you shot yourself.

Measure what that removes. A model sold for conversational video editing whose editing, where you live, only applies to what it made itself. For a videographer who wanted to retouch their own rushes, the tool does not do the thing they would have chosen it for.

No equivalent restriction is documented on the Wan 3.0 side.

This is the single criterion in this comparison where the right answer depends on your address rather than on the models. Same prices, same specs, two verdicts.


Our ranking, and a trap worth knowing

In the 9 September 2026 snapshot, the video category puts Wan 3.0 first with 99.46 and Gemini Omni 1.1 Flash second with 99.11. Both show 100 % measurement coverage and a verified status.

Thirty-five hundredths. Which is to say the question “which one is better” has no answer at that level of precision.

Two reading precautions, which we apply to our own data.

First precaution: the same model can appear twice. Wan 3.0 sits first with 99.46, fed by Artificial Analysis measurements — and a second time in third place with 97.56, under the spelling “Wan3.0”, without a space, fed by LMArena. Same model, two entries, two ranks. It is an aggregation artefact, and the rule that follows applies to anyone reading any leaderboard: check neighbouring spellings before quoting a rank.

Second precaution: two scores are only comparable if they come from the same family of sources. A score from LMArena and a score from Artificial Analysis measure things that are close but not identical. It is the same lesson we drew from setting Grok 4.6 against Gemini 3.7 Flash, where the same test gave 88.4 % in one version and 26 % in the next.


Audio: two approaches, one shared ceiling

Both models produce sound along with the image. Neither produces sound you can use as is for speech.

At Wan 3.0, sound is generated in the same pass as the image — dialogue, effects, ambience, music. The reports agree: ambience is good, footsteps, wind and crowd noise pass in a working edit. Dialogue, though, sounds processed. Alibaba’s announcement acknowledges this explicitly.

At Gemini Omni, the model tries by default to produce an appropriate audio track. But voice editing is not supported, audio references are not either, and the sound inside a reference video is ignored. Add the lip-sync drift the testers report.

Practical conclusion, valid on both sides: keep the ambience, redo the voice. Plan a separate audio track in your production chain.


On-screen text: the criterion English reverses

On paper the advantage was Chinese. In English it is not, and this is the one place where reading this comparison in English changes the answer rather than just the wording.

Wan 3.0 announces twelve supported languages for rendering text in the image. Gemini Omni’s spec sheet is more restrictive: English is fully supported, other languages have not been evaluated.

For a reader writing French, German or Spanish title cards, that restriction is a weakness and Alibaba has at least a chance. For an English caption it is the opposite: English is the only language either vendor claims to have evaluated, and it is Google that claims it.

Temper it immediately, as we do in every language. Alibaba itself writes that on-screen text accuracy is improving but not yet where it wants it, and testers report burnt-in text that is often wrong. Twelve announced languages are not twelve guaranteed languages.

And the gap is worth stating plainly: neither vendor publishes an evaluation in any language other than English. We flag the hole rather than filling it with an impression.


The watermark, and traceability

A structural gap that depends on no image quality.

Every Gemini Omni output carries the SynthID watermark — invisible to the eye, detectable programmatically. It is Google’s answer to the European obligation to mark synthetic content, in force since 2 August 2026, whose effects on video tools we detailed when comparing Kling 3.0 and Veo 3.1.

None of the sources consulted documents an equivalent watermark on Wan 3.0 outputs.

For a social media creator, the difference is invisible. For a newsroom, a school, a public body or a European legal department, it is decisive — and it leans towards Google, against the grain of the rest of this comparison. It also reaches further than Europe: anything you publish into an EU market is covered wherever you generated it.


What each admits about its own product

An exercise we run systematically: read what the vendor writes against itself.

Alibaba, in its official announcement of 13 August, writes that audio texture and on-screen text rendering accuracy are improving but not yet at the level it wants. That is in the launch press release, not in buried documentation.

Google acknowledges in its technical documentation that full consistency across edits, complex motion and perfectly accurate text remain difficulties. That is accurate and useful — but the public announcement leads with forty seconds that are a cumulative total and a 4K that is an upscale.

And each leaves a hole in the same place:

  • Google publishes no consistency measurement and no failure rate for scene extension, that is, for its main feature.
  • Alibaba cites no independent evaluation result in its announcement.

Those two absences are not oversights. They are the two figures that would decide the purchase.


Alibaba closed what it had opened

There is a reversal in this story, and it deserves saying.

The Wan family built its reputation on open weights. Its last version released under the Apache 2.0 licence is Wan 2.2. Since then, generations 2.5, 2.6, 2.7 and 3.0 are commercial models, available online only. No weights on public repositories, no self-hosting, no local GPU option: Wan 3.0 is driven from a browser or an API, and that is all.

So the vendor that carried the open-weights banner in video generation has closed its high end — exactly like the American vendors it competes with. It is the same movement we described around Kimi K3 and the licences attached to Chinese open weights: openness stays on the entry level, the top closes up.

If you want a model whose weights you can actually download and run on your own machine, that conversation now happens elsewhere — it is the one we had around Muse Glimmer and free ChatGPT, and it no longer includes the top of the video field.

What is still Chinese about Wan 3.0 is not openness. It is the price.


Which to choose

Take Wan 3.0 if you need a long shot in one take, if your raw material is a document — a deck, a product sheet, a web page —, if you work vertical for social media, or if you have to burn in text in a language other than English. It is also the default choice if you are in the United Kingdom or the European Economic Area and you want your own footage in the loop, because that is the one thing Google will not let you do there.

Take Gemini Omni 1.1 Flash if your work is made of back and forth — generate, look, fix a light, start again. Conversational driving with state preserved is real and comfortable, the draft mode at three cents makes exploration nearly free, English captions are the only ones either vendor says it evaluated, and the SynthID watermark settles a compliance question you will not have to raise. If you are in the United States, add the flagship feature back in: editing your own uploaded footage, which a British or Irish reader cannot use at all.

Take neither for speech. Dialogue sounds processed at Alibaba, lip-sync drifts at Google. Plan a separate audio track.

And take neither for a final delivery without an edit. Wan says so itself: it covers generation, not finishing. You will need editing software behind it, in both cases.


What to remember

  • Thirty-five hundredths of a point separate the two models in our ranking. Neither dominates.
  • Google’s forty seconds are a stack of three to ten second shots, appended at the end of a clip only, billed separately.
  • Its 4K is an upscale — and of four tiers, the two highest are, 1080p included.
  • Wan 3.0 generates thirty seconds in one pass, with sound: a continuous take, not an edit.
  • Editing your own video is blocked in the United Kingdom, the EEA and Switzerland at Google, at API level — and available in most of the United States. This is the one criterion where your country changes the answer.
  • Refusals are sometimes silent: zero tokens produced, empty output, no message.
  • At 720p the prices are identical. At 1080p Google is 25 % cheaper but interpolated, Alibaba dearer but native.
  • A delivered minute costs $14.40 against $21, attempts included — but one is made of six joins, the other of two takes.
  • Google’s real advantage is the three-cent draft, not the output price. This version matches Veo 3.1 Fast exactly.
  • On-screen text in English is Google’s, and that reverses the French edition’s conclusion: English is the only language either vendor says it evaluated.
  • Alibaba publishes its own list of weaknesses. Google publishes an announcement that survives its own documentation poorly.
  • Neither decisive gap is filled by anyone: no consistency measurement at Google, no independent evaluation at Alibaba.
  • The open-weights champion in video has closed its high end. The last published weights are Wan 2.2’s.