MiniMax-H3 — image-to-video prompt format
Twelve entries from the image_to_video prompt library, each run as
written and rewritten to MiniMax's I2VA format (integrated_multimodal_description /
overall_soundscape / non_diegetic_music, with the <Picture 1>
reference line). Everything else identical: 124 frames (5.17 s), 50 steps, seed 777,
SageAttention 2 + First Block Cache 0.10, RTX 5090 — generated by the
image-server port with comfy-kitchen absent. Source frames are Flux.2-Klein-4B.
Mean sampling 111s. 5 of 12 raw clips came out essentially silent (rms < 0.01) against 2 of 12 rewritten.
The rewrite is compute-free — measured, not assumed. All 24 clips ran
110.4–112.1s at 2.21–2.24s/it with 34/50 steps skipped; the ~1% spread is the longer text
encode. It costs encoder tokens, not sampling time.
Judge by watching and listening. ⚠️ Do not read the rms
figures as a quality score — several rewrites deliberately ask for near-silence ("almost
silent", "near silence"), so a low value there is the requested outcome, not a defect. What
the numbers do show is clips with no usable audio track at all: the raw prompts specify no
audio whatsoever and H3 synthesises a track regardless, so it fills the gap with whatever it
infers. The starkest pair is the steaming mug — raw 0.0014 (silent) against rewritten
0.2018.
The last four pairs test something separate — whether monochrome, sepia,
impressionist brushwork and flat art-deco graphics survive being animated, since a video
model's prior pulls hard toward photoreal. The rewrites restate the medium inside the motion
sentences rather than only at the top; the raw ones name it once.