Eight Seconds at a Time: What AI Video Generation Is Actually Good For
Generative video in August 2026 is a shot generator, not a film generator. The documented limits explain why, and they explain exactly where it earns its place.
Eight seconds. That is the honest working unit of AI video generation at the end of August 2026, and nearly everything interesting about the technology follows from that number.
Google's documentation for Veo 3.1 permits three durations: four seconds, six seconds or eight seconds. Eight is conditional, available only at 1080p or 4K or when you supply reference images. OpenAI's Sora guide is more generous on paper with 16 and 20 second generations, then adds a line the coverage skips. Characters, it says, work best in two to four second clips. Not sixteen. Two to four.
Showreels are cut from thousands of attempts. Documentation is written by engineers who have to support the thing. When the two disagree, believe the documentation.
Eight seconds is the real unit
Both major Western labs let you exceed a single generation. Neither lets you do it the way a director would want.
Veo 3.1 extends a clip by seven seconds at a time, up to twenty times, for a combined file of up to 148 seconds. Read the conditions. Extensions run at 720p only. The input must be a Veo-generated video no longer than 141 seconds. So 148 seconds is real, and it is a 720p continuous drift assembled seven seconds at a time, each pass conditioned on the last.
Sora works the same way with different arithmetic: 20 seconds per extension, six extensions, 120 seconds total. Runway's developer portal lists Seedance 2.5 at 1080p, cinematic video up to 30 seconds. That is the longest single-generation figure I can find in current vendor documentation, and it is a vendor figure.
None of this gives you a shot. A shot is a decision about where the camera sits and how long it stays there. These systems produce a continuous stretch of plausible motion, which is a different object, and the difference surfaces the moment you try to cut.

Content Credentials tell you a camera was there. They do not tell you whether anything true happened in front of it.
Consistency breaks at the cut
Inside one generation the current models hold together well. Faces stay faces. Water behaves like water. Hands have stopped being the joke.
Across two generations nothing carries over automatically. Not the jacket. Not the lens. Not the time of day, the wall colour or the geography of the room. Reference images and character features help, and every serious platform ships some version of them, but they nudge identity rather than specify it. Sora's guide caps you at two characters per video and says they hold best in two to four second clips, which is a fairly direct admission of where identity slips.
ByteDance states that Seedance natively supports narrative video across multiple cohesive shots with consistent subject, style and atmosphere. That is a vendor claim on a vendor page. The peer-reviewed literature is more careful. Mask-squared DiT (arXiv 2503.19881), presented at CVPR 2025, calls multi-scene video generation relatively underexplored, and its core attention scheme handles only a fixed number of scenes until a second conditional mask is bolted on to extend past them. When a conference paper still calls your production requirement underexplored, that is the more reliable signal.
Underneath the consistency problem sits the control problem. In an edit room you say: same room, camera thirty degrees left, two stops darker, hold on her hands. There is no interface for that instruction. You re-prompt, you re-roll, and the system hands you a different room.
Where AI video generation earns its place today
Four uses hold up because of the constraints above rather than in spite of them.
Previsualisation. Previs is disposable by design. Nobody grades it, nobody delivers it, and the point is to be wrong quickly. Eight seconds is roughly the length of a previs beat anyway.
Concept and pitch work. When you are arguing for a direction, the room needs to see the idea, not the final frames. A generated mood reel does what a stack of stills used to do, with motion.
Short-form marketing. Vertical social video is natively short, natively fast and tolerant of a slightly synthetic look. It is the one format where the eight second ceiling binds nothing, because the format's own ceiling is lower.
Iteration before an expensive shoot. This is the strongest case and the least discussed. The scarce resource on a production is never compute. It is a crew standing on a location with the light going. Anything that moves a decision from the shoot day to the week before earns back the generation cost many times over, and generated tests are a cheap way to argue about blocking, framing or wardrobe while everyone is still at a desk.
What does not hold up: anything with a named person who must look identical in shot twelve and shot forty. Anything where the product has to be the product. Anything with continuity a viewer will check.
Cost per keeper
List prices, current at the time of writing. Veo 3.1 is $0.40 per second at 720p and 1080p, $0.60 at 4K, audio included by default. Veo 3.1 Fast is $0.10 at 720p and $0.12 at 1080p. Sora 2 is $0.10 per second at 720p; sora-2-pro is $0.30 at 720p and $0.70 at 1080p, with fifty percent off in batch. Runway's Gen-4.5 works out near $0.12 per second on the developer API by my own reckoning in the Runway review on this site, with Gen-4 Turbo around $0.05.
An eight second Veo shot with sound therefore costs $3.20. That is the number people quote.
Here is the number nobody quotes. You do not get one keeper per generation. On a brief with any specificity, several attempts per usable clip is the honest planning assumption, which puts a ten second finished Gen-4.5 clip closer to five dollars than to the nominal $1.20. Build thirty seconds from four shots and each pass generates thirty-two seconds of output, roughly thirteen dollars on Veo, and you will not accept the first pass. Six passes still leaves you under a hundred dollars of compute for thirty finished seconds. Genuinely cheap.
The cost that did not move is the human one. Somebody wrote the shot list, judged every take against the one before it, re-prompted, graded, cut and mixed. At any professional rate that labour dominates the compute by an order of magnitude, and it dominates it more as compute gets cheaper.
I wanted to set a conventional production figure beside this. I am not going to. The averages that circulate for commercial production cost trace to trade association surveys behind membership paywalls, or to nothing at all, and I could not open a primary document for any of them. A number I cannot check is worse than no number. The structural claim stands without one: generated video drives the marginal cost of one more take close to zero and leaves the cost of deciding which take was right exactly where it was.
Content Credentials prove capture, not authorship
The provenance stack is further along than most people assume and answers a narrower question than most people want. C2PA is at specification version 2.3, steered by a committee including Adobe, Amazon, the BBC, Google, Meta, Microsoft, OpenAI, Publicis Groupe, Sony, TikTok and Truepic. That is most of the capture chain and most of the distribution chain in one room, which is why the standard will probably stick.
Now read what the specification says about itself. C2PA specifications should not provide value judgments about whether provenance data is good or bad, only whether the assertions can be validated as associated with the asset, correctly formed and free from tampering.
That is the whole thing in a sentence. A Content Credential proves a set of claims was cryptographically bound to this file and has not been altered since. It proves capture in the narrow sense that a signing device asserts this file came off it. It does not prove authorship and it does not prove truth. A camera with a signing chip will sign a staged scene without complaint. Every frame verifiably unmodified. Every frame verifiably false.
Google's SynthID takes the other route, embedding an invisible watermark at generation time that DeepMind says survives cropping, filters, frame rate changes and lossy compression. Detection runs through Google's own tooling, which is useful and is also a vendor-controlled channel.
The practical failure is duller than either. Metadata does not survive the internet. Re-encode, re-upload or screen-record and the manifest is gone while the video is fine. A standard that validates perfectly in a lab and strips on the second share is a standard about institutions rather than about files.
Four of the top five video models are Chinese
Artificial Analysis runs a Video Arena that ranks models by Elo from blind human preference votes rather than vendor benchmarks. On its text to video board at the time of writing, the top five are Alibaba's Wan 3.0 on 1,242, Google's Gemini Omni Flash on 1,237, MiniMax H3 Max as post-trained by fal on 1,235, MiniMax H3 on 1,227 and Dreamina Seedance 2.0 at 720p on 1,221. Veo 3.1 sits thirteenth on 1,092. Sora is absent from the board.
Two honest caveats. Preference on short prompts measures immediate appeal, not usability across a sequence, and a model that wins an aesthetic A/B can still fall apart at the fourth cut. Sample sizes vary widely, from around five thousand votes to nearly twenty-five thousand.
What matters more than the ranking is the licensing. Wan 2.2 shipped under Apache 2.0 with no field of use restriction, no monthly active user cap and no territorial carve-out. From Hangzhou that reads as ordinary. It is not. It means part of the current frontier is available as weights you can run yourself, and the cost floor for a studio in Lagos or Nairobi becomes a GPU rather than a recurring bill in dollars.
The camera was never the bottleneck
Here is the argument against everything above, and it is a good one.
The constraint on video was never equipment. A phone has shot broadcast-usable images for the better part of a decade. Editing software has been free and excellent for longer. If cheap capture produced good work, the last fifteen years would have buried us in masterpieces.
What was scarce was having something worth saying and the craft to shape it. A script that earns its runtime. A performance. An editor who knows the cut lands three frames earlier than it feels. None of that is generated, not by these models and not by the next ones.
Push it harder. When competent-looking footage becomes effectively unlimited, the scarce good stops being footage and becomes attention, and generated video worsens that oversupply for everybody including the people producing it. The advantage goes to whoever can already tell which of six takes was right. That person usually learned it on sets, which is to say they already had access to budgets.
This is not a small objection. It is most of the objection, and about outcomes I think it is correct.
Budget was a constraint, not the constraint
Where I part company is on the word bottleneck, which quietly assumes there is only one.
In markets where a location permit, an insurance bond, equipment rental and twelve day rates set the entry price, that price filtered who got to attempt anything at all. Removing it does not create talent. It changes the size of the population that gets to find out whether it has any. Those are different claims and only the second is being made here.
I would like to size that population. The figures everyone cites for Nigerian film output and for African audiovisual GDP trace to UNESCO's 2021 report The African Film Industry, which is open access at unesdoc.unesco.org, so anyone can check them at source rather than take them from the coverage. What I can describe is the friction as it appears from here, which is not the friction the coverage describes.
It is payment rails. A large share of these APIs will not take a Nigerian card, which makes list pricing academic. It is bandwidth, because iterating on video means moving video, repeatedly, over connections metered by the gigabyte. It is currency, because forty cents a second is not forty cents a second when you earn in naira.
And it is the training data. Ask any of these models for a specific street in Lagos, a specific market in Kano, a specific fabric with a specific pattern. You get an average of everything labelled African that the model has seen. Confident. Well lit. Wrong in a way that is hard to point at.
So the tool that finally makes production affordable for these markets is also the one worst at rendering them accurately, and its output will carry a Content Credential attesting, correctly, that nothing was tampered with. The cost of making a picture of a place has collapsed. The cost of making a picture of a place that does not exist collapsed by exactly the same amount. For regions that have spent a century being described by other people, that is a strange thing to be handed cheaply.
Tools referenced
Runway, reviewed here: Runway review.
Google Veo 3.1, reviewed here: Google Veo 3.1 review.
Kling AI / 可灵, reviewed here: Kling AI / 可灵 review.
Wan 2.2, reviewed here: Wan 2.2 review.
Hailuo AI / 海螺 AI, reviewed here: Hailuo AI / 海螺 AI review.
Descript, reviewed here: Descript review.
Sources
OpenAI, Video generation API guide (sora-2 and sora-2-pro durations, extension limits, character limits): https://developers.openai.com/api/docs/guides/video-generation
OpenAI, API pricing (per-second Sora rates by resolution, batch discount): https://developers.openai.com/api/docs/pricing
Google, Gemini API Veo documentation (Veo 3.1 durations, resolutions, extension mechanics and storage): https://ai.google.dev/gemini-api/docs/veo
Google, Gemini API pricing (Veo 3.1 and Veo 3.1 Fast per-second rates, audio included): https://ai.google.dev/gemini-api/docs/pricing
Artificial Analysis, Video Arena text-to-video leaderboard (Elo rankings and vote counts): https://artificialanalysis.ai/video/leaderboard/text-to-video
C2PA Specification 2.3: assertions, validation and tamper detection: https://spec.c2pa.org/specifications/specifications/2.3/specs/C2PA_Specification.html
Coalition for Content Provenance and Authenticity (specification version 2.3, steering committee membership): https://c2pa.org/
Google DeepMind, SynthID (video watermarking and stated resilience to crops, filters, frame rate and compression): https://deepmind.google/science/synthid/
Frequently Asked Questions
How long can AI video generation produce in a single clip in 2026?
Google's documentation for Veo 3.1 permits four, six or eight seconds, with eight available only at 1080p or 4K or when reference images are supplied. OpenAI's Sora guide offers 16 and 20 second generations. Runway's developer portal lists Seedance 2.5 at up to 30 seconds of 1080p. Anything longer comes from chained extensions rather than single generations, and on Veo those extensions run at 720p only, up to a combined 148 seconds.
Why do generated characters change between shots?
Because nothing carries between generations by default. Each generation is a fresh sample conditioned on your prompt and any reference images you supply, and reference conditioning nudges identity rather than specifying it. OpenAI's own Sora guide caps you at two characters per video and notes they hold best in two to four second clips, which tells you where identity begins to drift.
Do C2PA Content Credentials prove a video is real?
No. The C2PA specification is explicit that it validates whether assertions are correctly formed, bound to the asset and free from tampering, and that it should not judge whether provenance data is good or bad. Credentials establish that a file has not been altered since signing. They say nothing about whether the scene in front of the lens was staged, and they establish nothing about authorship.
Is generative video actually cheaper than filming?
Per second of output, dramatically. Veo 3.1 lists at $0.40 a second with audio and Sora 2 at $0.10 a second at 720p. The saving shrinks once you count the takes discarded per usable clip, and it stops mattering once you count the human hours spent writing the shot list, judging takes and cutting. Compute falls. Judgement does not.
Which video generation models rank highest right now?
On Artificial Analysis's Video Arena, which ranks by blind human preference rather than vendor benchmark, the top five text to video entries at the time of writing were Alibaba's Wan 3.0, Google's Gemini Omni Flash, MiniMax H3 Max, MiniMax H3 and Dreamina Seedance 2.0, with Veo 3.1 thirteenth. Preference votes on short prompts measure immediate appeal rather than usability across a sequence, so read the board as a signal about look, not about production fit.
Read next
China's University Major Cuts Are AI Policy, and Nigeria Should Read the Fine Print
The latest analysis essay.
Working on something in this space?
If this analysis is close to a problem you're thinking about, say so. I read every message personally.
Start a conversation