One prompt per category from a public prompt library, run as published and rewritten to MiniMax's T2VA format. Everything else identical: 1344×768, 243 frames (10.1 s), 50 steps, seed 12345, SageAttention 2 + First Block Cache + mmgp block offload, one RTX 5090. The rewrite is compute-free — it costs encoder tokens, not sampling time. Judge by watching and listening: the original prompts specify no audio at all, and H3 generates audio either way.
Generation cost is set by resolution, frame count and steps — not by the prompt, so it is reported once rather than per variant. First Block Cache makes steps bimodal, so a single seconds-per-iteration figure cannot be extrapolated: at this threshold it skips ~68% of steps, and a skipped step costs 0.39s against 18.3s for a full one.
| full | cached | generate | without FBCache | |
|---|---|---|---|---|
| 20 steps | 6 | 14 | 152s | 402s |
| 25 steps | 8 | 17 | 190s | 494s |
| 30 steps | 10 | 20 | 227s | 585s |
| 50 steps | 16 | 34 | 342s | 950s |
social media