Every AI video provider publishes a price per second. Almost no finance team should budget from it.
In short: the published rate covers one render. Production budgets are set by the number of renders it takes to get a usable one — a multiplier that varies more between models than the sticker price does, and that nobody publishes.
The per-second rate is the easiest number to compare, which is why procurement conversations start there and often end there. Current API-tier pricing runs from roughly $0.08 per second at 768p on Hailuo’s MiniMax H3 to around $0.19–$0.32 per second for ByteDance’s Seedance 2.5 at 720p, depending on whether a reference video is supplied. On a spreadsheet that looks like a four-fold difference, and the decision appears to make itself.
Then the first real project lands and the spreadsheet stops predicting the invoice.
What the sticker price leaves out
A finished eight-second clip is not one render. It is the render that shipped, plus every render before it that didn’t. Teams that track this properly tend to find the ratio sits somewhere between three and ten attempts per usable output, depending on how specific the brief is and how well the model responds to correction.
That multiplier does the real work in the budget. A model at $0.08 per second that needs eight attempts costs more per delivered clip than a model at $0.19 per second that lands in three. The published rates say the opposite, which is why the published rates mislead.
The multiplier is also the part no vendor publishes, because it depends on your prompts, your subject matter, and your tolerance for “close enough.” It has to be measured in-house, and it can be measured cheaply — the cost of running a controlled comparison for an afternoon is trivial next to the cost of standardising on the wrong model for a year.
Three things that move the multiplier
Whether a single prompt edit produces a single change. This is the most useful property to test and the least discussed. Change one clause in the prompt and re-run. If the whole scene reorganises — different framing, different lighting, a different face — you cannot iterate, you can only re-roll. Models that respond proportionally to a small edit let a reviewer converge in two or three passes. Models that don’t turn every note into a fresh gamble, and the attempt count climbs accordingly.
Whether the failure is legible. When an output misses, some models miss in a way you can diagnose — the prompt asked for two things and the model prioritised one. Others miss opaquely. Diagnosable failures cost one more attempt. Opaque ones cost a session.
Who is allowed to press the button. If regeneration requires a specialist with a licensed seat, iteration is capped by that person’s calendar rather than by the model. Teams routinely evaluate on output quality, standardise on the best single result, and then rediscover the queue they were trying to remove. Widening access usually buys more throughput than upgrading the model does.
How to run the comparison in an afternoon
Take three briefs that resemble your actual work — not showreel material, the boring corporate case. For each model, run the brief, apply one written note, re-run, and repeat until the output would pass internal review. Record two numbers: attempts to acceptance, and total seconds rendered.
Cost per delivered clip falls straight out of that, and it is the only number worth putting in a budget. Three briefs across two or three candidate models is typically under twenty dollars of compute and gives you a defensible figure instead of an inherited assumption.
A few practical notes on running it fairly. Use the same prompt text across models rather than tuning each one — you are measuring how the model handles imperfect instruction, which is the condition it will actually work under. Have the same person judge acceptance for every run. And include at least one brief with a specific constraint (a product colour, a piece of on-screen text, a required camera move), because that is where attempt counts diverge most.
Where the per-second rate does matter
Two cases. The first is high-volume, low-specificity work — background plates, ambient loops, filler where any plausible output passes. There the multiplier collapses toward one and the sticker price becomes the whole story.
The second is the ceiling. If a model cannot produce an acceptable result at any attempt count, its rate is irrelevant. Resolution caps, maximum clip length, whether audio is generated natively, and whether the model accepts a reference video are all pass/fail gates that should be checked before any pricing comparison begins. There is no point optimising the cost of a model that cannot do the job.
Between those two poles — which is where most commercial work sits — the multiplier dominates, and it is measurable in an afternoon.
The version of this that gets budget approved
Rather than “Model A is $0.08 and Model B is $0.19,” the sentence that survives scrutiny is “Model A delivers an accepted clip for $2.40, Model B for $1.90, measured across three representative briefs on this date.” It is a harder sentence to produce and a much harder one to argue with.
It also ages honestly. Rates change, models get replaced, and a rate-based decision has to be redone from scratch each time. A measurement harness gets rerun in an afternoon.
Teams comparing current-generation video models can find per-model specifications, published API rates and worked cost examples on Kavel, including a breakdown of MiniMax H3 resolution tiers and per-second pricing.



