Key points
- Single-pass clip length is a hard architectural constraint, not a product tier. Coverage as of August 2026 runs roughly 8–30 seconds depending on the model, with longer runtimes produced by chained extensions.
- OpenAI's published API for Sora exposes discrete clip durations — 4, 8 and 12 seconds, with longer values available on current documentation — rather than an arbitrary length parameter. Discrete options are a sign of how the models are built.
- Extension is a splice, not continuity. Each chained segment inherits an approximation of the previous state, so characters, lighting and lens behaviour drift at every seam. Drift compounds; it does not average out.
- The professional workaround is shot-based production: write to shots of a few seconds, generate each independently, and assemble. This is how film has always worked, and it removes the constraint from the critical path.
- Continuity is bought with references, not prompts. Locked character references, a written design bible, fixed lens and lighting language, and a plate-and-composite approach for anything that must match exactly.
- The limit is not the binding constraint on most enterprise video. Corporate explainers, product walkthroughs and ads are already cut from shots of two to six seconds. The constraint bites on unbroken long takes, which are rare in commercial work.
Disclosure. Published by Lifewood Data Technology, which produces AI-generated video commercially. Model capabilities in this area change on a timescale of months; figures are given with dates attached and should be re-checked against the current documentation of any model you are about to build a pipeline around. Cross-model comparison figures come from a secondary industry survey and are labelled as such.
Why the limit exists, and why it is not simply a pricing decision
It is tempting to read clip-length caps as artificial — a tier boundary that a bigger plan would remove. They are not, mostly. A generative video model has to produce every frame in a way that is consistent with every other frame: the same face, the same jacket, the same room, lit the same way, obeying the same physics. The computational cost of maintaining that consistency does not grow linearly with duration, and neither does the number of ways it can fail. Doubling the length more than doubles the chance that a hand becomes six fingers halfway through, or that a background sign quietly rewrites itself.
That is why published interfaces tend to offer discrete durations rather than a free-form length field. OpenAI's video generation API for Sora exposes a small set of allowed clip durations rather than an arbitrary seconds parameter — a design that reflects how the generation is actually structured. The same pattern shows up across the field.
As of August 2026, an industry survey of published limits put the single-pass range across major models at roughly eight seconds at the conservative end to about thirty seconds at the longest, with several widely used models clustered around fifteen. Longer runtimes are advertised, but reached by chaining: the model extends an existing clip by generating a continuation conditioned on its ending. Those are different products, and the difference is visible on screen.
What chaining actually costs you
Chained extension is the obvious workaround and it works, within limits worth understanding before a production depends on it. The model generates a continuation conditioned on the end of the previous segment, which gives approximate rather than exact continuity. The consequence is drift, and drift has a characteristic signature.
- Identity drift. A face or a costume detail shifts slightly across a seam. Individually imperceptible; across six seams, the character at the end is visibly not the character at the start.
- Lighting and grade drift. Colour temperature and contrast wander, which is the easiest artefact to spot in a finished cut and, fortunately, the easiest to correct in the grade.
- Physics and motion discontinuity. Momentum is not conserved across a seam, so a moving object can change speed or direction at the join.
- Background instability. Text on signage, patterns and crowd detail regenerate rather than persist, because they were never represented as stable objects.
The important property is that these accumulate. Each seam adds a small error and the next segment inherits the drifted state as its reference, so a long chain diverges monotonically. This is why no model advertises unlimited chaining even where the documentation sets no explicit cap — the useful limit is set by acceptable drift rather than by the API.
Editorially, none of this is unusual. Traditional production also cannot hold a shot forever, and it solved the problem the same way: cut. The mistake is treating chaining as a way to avoid learning shot-based construction, rather than as one tool inside it.
Shot-based production: the architecture that removes the constraint
This is the workflow that turns a seconds-long generation ceiling into a non-issue. It is deliberately close to conventional film production, because conventional film production is a solved answer to the same constraint.
Producing long-form video from short generations
-
Write to shots, not to scenes
Break the script into shots of two to six seconds with a stated purpose for each. If a beat cannot be expressed in shots of that length, that is a writing problem worth solving before generation — long unbroken takes are rare in commercial video for reasons that predate AI.
-
Lock a design bible before generating anything
Character references, wardrobe, location plates, lens language, colour palette, lighting direction, and grade target — written down and, wherever the tooling supports it, stored as reference images. Continuity is bought here, not recovered later.
-
Generate each shot independently against the bible
Every shot references the same locked assets rather than the previous shot's output. Independent generation means a failed shot is one regeneration, not a re-run of the whole chain — which is also what makes iteration affordable.
-
Over-generate deliberately and select
Produce several takes per shot and choose in the edit. Generation cost per take is low relative to review cost, so the economics favour selection over prompt refinement; this is the closest analogue to shooting coverage.
-
Reserve chaining for genuine long takes
Where a beat truly requires unbroken motion, chain — and budget grade and cleanup work at each seam. Treat it as an effects shot with a known cost, not as the default construction method.
-
Composite anything that must match exactly
Logos, product renders, UI screens, legal text and precise brand colour should be composited over generated plates rather than generated. Models approximate; brand assets cannot be approximated.
-
Assemble, grade and sound design conventionally
Cut in an editor, grade to a single target to absorb residual drift, and build the sound design and music as one continuous layer across the shots. A consistent audio bed is remarkably effective at binding visually heterogeneous shots into one piece.
-
Deliver with provenance and labels attached
Mark the output as AI-generated at export, apply any required visible disclosure, and record model, version and date per shot. Retrofitting this across a shot-based project after delivery is considerably harder than doing it at export.
Where the constraint actually bites
Whether clip length matters depends almost entirely on what is being made, and for a large share of commercial video the answer is that it does not.
| Format | Constrained? | Why |
|---|---|---|
| Social and short-form ads | Barely | Already cut from shots of one to three seconds; the ceiling is above the shot length |
| Product explainers and walkthroughs | Barely | B-roll over narration; shots are short and the voice track carries continuity |
| Localized variants of an approved master | No | Picture is locked; the variable is language, not shot length |
| Presenter-led corporate video | Somewhat | Sustained on-camera performance is exactly the case where seams show; consider a real presenter with AI-generated surrounds |
| Narrative film and drama | Yes | Performance continuity across long takes is the hardest thing to hold, and the thing audiences are most sensitive to |
| Documentary with archival material | Yes, differently | The constraint is provenance and permissibility rather than duration — synthetic reconstruction of real events raises disclosure duties |
The row worth planning around is presenter-led video. A hybrid — a real presenter shot conventionally against generated environments and B-roll — usually beats a fully synthetic presenter for anything longer than about thirty seconds, and it sidesteps the likeness-consent questions covered in the rights guide.
What this means for catalogue-scale production
The strategic point is that shot-based construction is not a workaround. It is the same architecture that makes multilingual variants cheap, described in the localization guide: a master built from separable components, where each component can be regenerated or swapped without touching the others. A production built as one long generated take is as brittle as a video with burned-in subtitles — any change means starting over.
Built the other way, a catalogue of hundreds of videos becomes tractable: a shared design bible, a shot library that can be reused across titles, per-title generation only for the shots that are genuinely specific, and language variants layered on top of a locked picture. The clip-length ceiling stops being a limitation and becomes a unit of work — which, in a production system, is what you want a limitation to turn into.
Lifewood Data Technology produces AIGC video at catalogue volume on this architecture, with human creative direction at the shot level and region-native review on every language variant. A current example of the scale this is built for: a publishing engagement signed 29 April 2026, a two-year framework of approximately USD 3 million covering up to 3,000 titles, with two roughly 45-second trailers and one roughly 3-minute promotional video per selected title. Nothing in that brief requires a single generation longer than a few seconds.
Related questions
What is the longest video an AI model can generate in one pass?
As of August 2026, published single-pass limits across major models run from around eight seconds at the short end to roughly thirty at the long end, with many clustered near fifteen. Longer advertised runtimes are produced by chaining extensions rather than by a single generation. Because this changes on a monthly cadence, check the specific model's current API documentation rather than any comparison article.
Why can't I just chain clips to make a five-minute video?
You can, and the result will drift. Each extension conditions on the previous segment's ending and reproduces it approximately, so identity, lighting, motion and background detail shift slightly at each seam and the error accumulates down the chain. For a five-minute piece, shot-based construction with conventional editing produces a materially better result at lower cost.
How do you keep a character consistent across separately generated shots?
With locked reference assets rather than with prompt wording. Fixed character reference images, a written wardrobe and design bible, consistent lens and lighting language, and a single grade target applied in post. Where exact matching is required — a logo, a product, a UI screen — composite the real asset over a generated plate instead of asking the model to reproduce it.
Is AI video suitable for a presenter-led corporate film?
Partly. A fully synthetic presenter holds up for short durations and starts to show seams over longer sustained performance, and it raises likeness and consent questions covered in the rights guide. The pattern that works well today is hybrid: a real presenter shot conventionally, with generated environments, B-roll and graphics around them.
Does the clip-length limit affect localized versions of a video?
No. Localization operates on a locked picture — the variables are script, voice, on-screen text and subtitles, none of which involve regeneration. This is one of the reasons the master-and-layers architecture is worth building even for a project that is currently single-language.
Should we wait for models that generate longer clips?
Waiting optimizes for a constraint that mostly does not bind. The formats where clip length actually limits what can be made — narrative drama, long unbroken takes — are a small share of commercial video, and the architecture that solves the problem today is the same architecture that will make longer models useful when they arrive. Building shot-based now is not a stopgap.
Sources
Every figure, date and legal requirement above traces to one of these. Where a source is a market-research summary, a vendor pricing page or a preprint rather than a primary document, the text says so at the point of use.
- Video generation with Sora — API guide, including supported clip durations
- Create video — API reference for the videos resource
- Veo — model overview
- How Long Can AI Videos Be? Maximum length by model, August 2026
- Article 50 — Transparency Obligations for Providers and Deployers of Certain AI Systems
Talk to the Lifewood AIGC team
Lifewood builds AIGC video shot-first, with a locked design bible per programme and human creative direction at the shot level, which is what makes catalogue-scale output and fifty-language variants the same pipeline rather than two. Send a script and a shot count and the response includes the shot breakdown, not just a price.