Video
Generate a video clip from a description
A sentence in, a few seconds of video out. Which one, and what does a minute of finished footage actually cost?
Last checked 2026-08-29How we ranked this ↓
1
Gemini Omni Flash
Google#1 in blind tests (no audio)Top of the blind arena, ten cents a second, and you can edit it by talking to it.
- Use it when
- You want the best-looking short clip for the least money, and you expect to revise it — 'same shot, slower push in' rather than starting the prompt over.
- Cost
- About $0.10 per second of 720p.About $0.10 per second of 720p video on the Gemini API (billed per token — 5,792 tokens per second of 720p at $17.50 per million output tokens). Also in the Gemini app and Google Flow.
- Good at
- Conversational editing and multimodal input — it takes text, images and video together and keeps text and action in sync. Launched into public preview on 30 June 2026 and went straight to the top of the no-audio arena on Elo 1324, above every model here.
- The catch
- It is a public preview, and Google publishes its own list of what doesn't work yet: no audio references, no scene extension through the API, no reliable video references, and character consistency degrades across scene changes. It also caps at ten seconds. Preview also means the behaviour and the price can move without your project's permission.
- Wrong for
- Anything longer than ten seconds in one pass, anything where a character has to survive a cut, and anything on a deadline you can't afford to have a preview model break.
2
Veo 3.1
GoogleThe shipped one. Native audio, real 4K, and a proper editor around it in Flow.
- Use it when
- You need something finished rather than something impressive — audio included, longer than a preview model allows, inside a tool built for sequencing shots.
- Cost
- $0.05–$0.60 per second, depending on tier.Per second on the Gemini API: Lite $0.05 (720p) / $0.08 (1080p), no 4K. Fast $0.10 (720p) / $0.12 (1080p) / $0.30 (4K). Standard $0.40 (720p and 1080p) / $0.60 (4K). Also included in Google AI subscription tiers as Flow credits — check current pricing for those.
- Good at
- Being a product rather than an endpoint. Flow gives you shot sequencing, extension and reference handling, and Veo has been in the field long enough that the failure modes are documented by other people rather than discovered by you.
- The catch
- There is a twelve-fold price spread between Lite 720p and Standard 4K, and most people try Lite first, decide AI video looks cheap, and never learn that the version they were shown in the demo cost eight times more per second. Pick the tier deliberately. Google's own newer model also outranks it in blind votes at a quarter of Standard's price, so you are buying maturity here, not the frontier.
- Wrong for
- Bulk experimentation at Standard tier — the per-second cost punishes iteration, which is exactly what early drafts need.
3
Runway
RunwayThe editor, not the model. It runs its own Gen-4.5 and everyone else's models in one place.
- Use it when
- The generation is one step in a longer process — video-to-video, motion control, inpainting, ProRes delivery — and you'd rather not stitch five vendors together.
- Cost
- $12–$76/mo. The $12 plan buys ~52 seconds of video.Free: 125 one-time credits. Standard $12/month billed yearly, 625 credits/month. Pro $28/month yearly, 2,250 credits. Max $76/month yearly, 9,500 credits. Gen-4.5 video costs 60 credits per 5 seconds. Paid plans include Gen-4.5, Gen-4, Act-Two, Aleph and third-party models including Veo 3.1 and Kling 3.0 Pro.
- Good at
- Control. Video-to-video, motion controls, professional delivery formats and workflows — the parts of the job that a raw model API leaves entirely to you.
- The catch
- Do the credit arithmetic before subscribing. 625 credits at 60 per 5 seconds is about 52 seconds of Gen-4.5 a month on the $12 plan — under a minute of footage, and every rejected take spends from the same pot. Credits don't roll over on Standard or Pro (Max keeps up to a month). The free tier's 125 credits are one-time, not monthly, so it's a demo rather than a trial.
- Wrong for
- Volume generation. Per second of finished video this is among the most expensive routes here, and you're paying for the editor whether you open it or not.
4
MiniMax H3
MiniMaxSecond in the blind arena, with published weights — and a licence you must actually read.
- Use it when
- You want frontier-adjacent quality with native synchronised audio, and either self-hosting or a non-US/EU jurisdiction makes an open-weight model genuinely useful to you.
- Cost
- Weights free. Hosted pricing not published.Weights are free to download under the MiniMax H3 Community License. Hosted API pricing through MiniMax — check current pricing, we could not open a published rate card.
- Good at
- One-pass audio-visual generation: 33B parameters, text, image, video or audio in, 4 to 15 seconds out at 24fps with 32kHz stereo sound, up to 2K through a separate regeneration module. Elo 1301 without audio, second only to Google.
- The catch
- Two problems, both easy to miss. The licence: the model card links a separate application form specifically for the USA, EU, UK and South Korea, and secondary readings of the community licence describe those regions as excluded from default local deployment rights — read the licence yourself before you deploy, because 'the weights are on Hugging Face' is not the same as 'you may use them'. And the download isn't the whole model: the instruction-refinement layer is hosted rather than open-sourced, and the module that lifts output to 2K is not open-sourced either, so what you can run at home is not what you voted for in the arena.
- Wrong for
- A US or EU company that wants to self-host without legal review, and anyone without serious GPU capacity — 33B dense is not a laptop model.
5
Kling 3.0
KuaishouThe cheap one that people keep choosing anyway, for multi-shot consistency.
- Use it when
- You need several shots of the same subject to look like the same subject, and the budget is the binding constraint.
- Cost
- Couldn't verify — their pricing pages wouldn't load.Sold as subscription credits with per-second API rates. Check current pricing — Kling's own pricing and API pages would not load for us on 2026-08-29, so we are not quoting a number we could not read at source.
- Good at
- Subject and scene consistency across multiple shots at a price well below the Western frontier tiers, which is why it keeps turning up in production workflows despite the friction.
- The catch
- We could not verify its pricing from Kling's own site, which is itself the story: it is the least legible tool in this list from outside China. Subscription credits are reported to expire monthly rather than accumulate, content moderation is stricter and less predictable than the Western tools, and if you have data-residency obligations, a Kuaishou-operated endpoint is a conversation with your legal team, not a checkbox.
- Wrong for
- Regulated industries, anything with a compliance review, and anyone who needs to forecast cost precisely before committing.
Also considered
What we left out, and why. A list is only trustworthy if you can see what it rejected.
- Sora 2 / Sora 2 Pro — Excluded on a hard date, not an opinion. OpenAI's own deprecation page lists sora-2, sora-2-pro, their snapshots and the entire Videos API for removal on 24 September 2026, with no replacement named — under a month from this review. The consumer app went earlier. Do not start anything on it.
- Wan 3.0 — Alibaba's model leads the with-audio arena on Elo 1241 and is reported to do 30 seconds in a single pass from text, image, audio, video or documents. Left out because we could not open Alibaba's own pages to verify anything, access is a beta you apply for on Model Studio, and the open-weights story is genuinely contested — sources disagree about whether 3.0 weights exist at all or whether the open line stops at 2.2. Promising and unverifiable is not a recommendation.
- Dreamina Seedance 2.0 — ByteDance's model sits in the top five of both arena boards, which is a real result. Left out of the ranking for the same reason as Wan: we could not confirm access terms or pricing from the source.
- HappyHorse 1.0 / 1.1 — Third and fifth on the no-audio board. Two research-stage models from Alibaba with no product around them that we could find. Watch the leaderboard, not this entry, for these.
- Luma, Pika and the rest of the 2024 cohort — Still running, still cheap, and comprehensively out-ranked. If you already pay for one it is fine; there is no reason to start now.
What would change this list
- Gemini Omni Flash leaving public preview. It already tops the arena at $0.10 a second — the only thing keeping it from being an unqualified recommendation is that Google reserves the right to change it.
- 24 September 2026: the Sora 2 API and OpenAI's Videos API shut down. Anything built on them stops working, and OpenAI has not named a successor.
- Whether Alibaba ships open Wan 3.0 weights under Apache 2.0 as some sources claim is planned. An openly licensed model at the top of the with-audio board would reorder this entire list.
- The two arena boards converging. Right now audio and no-audio rank differently, which means nobody has yet built the model that is best at both.
How we ranked this
Sources
Artificial Analysis — text-to-video arena (both boards)Google — Gemini API pricing, per-second video ratesGoogle — Gemini Omni Flash launch and stated limitationsOpenAI — deprecations page, Sora 2 and Videos API shutdownRunway — pricing, credits and included modelsMiniMax H3 — model card and licence