MiniMax H3

MiniMax H3 Is Here: A Video Model That Also Edits

Share This Spread Love
Rate this post

Image generation went through this exact arc, and it’s worth remembering how it ended.

The first wave was pure text-to-image. You wrote a prompt, got a picture, and if it was nearly right you wrote another prompt and got a different picture. People built elaborate folklore around prompt phrasing because prompting was the only lever available.

Then editing arrived — masking a region, changing one element, extending a canvas, removing an object — and the folklore evaporated almost overnight. Not because generation got worse, but because nobody actually wants a new image. They want their image with a fix in it. Within a year the editing tools were what professionals used daily and pure generation was where you started, not where you worked.

Video has been stuck in the first phase considerably longer, for the obvious reason that preserving a scene across hundreds of frames is much harder than preserving a still. MiniMax H3 is a fairly clear marker that the second phase is beginning.

Why the second phase matters more

The reason the arc repeats is that generation and editing answer different questions.

Generation answers can this model produce something good. Editing answers can this model produce the specific thing I need, and then change one part of it without ruining the rest. Professional work runs almost entirely on the second question, because almost all professional work involves someone asking for a change to something that already exists and has already been approved.

The technical demand is different too, and harder. Generating requires plausibility. Editing requires plausibility plus preservation — alter the designated element while every other element stays exactly as it was, frame after frame: the same face, the same light direction, the same camera path, the same background action at the same moments.

A model that regenerates the scene with your change applied hasn’t edited anything. It’s made a second video resembling the first, and every approval you’d already won is back in play.

What H3 actually does to footage

Minimax H3 modifies existing video locally. The supported operations map closely onto what revision requests actually consist of:

  • Replacing, adding, or removing objects and people
  • Changing backgrounds, environments, and lighting
  • Adding or adjusting visual effects
  • Modifying a character’s motion or performance
  • Replacing dialogue, with the mouth re-forming around the new line
  • Cloning or transferring vocal timbre

It currently leads the Artificial Analysis video editing leaderboard, ahead of Seedance 2.0 — a ranking that measures preservation as much as fidelity, which is precisely the property this kind of work depends on.

That dialogue capability is worth pausing on, because it’s the one with no counterpart in image editing. Changing a spoken line means changing picture and sound together, in sync, with a performance that still reads as one continuous take. It’s only possible because audio isn’t a separate stage here.

The audio half

Every H3 output carries native stereo audio generated in the same pass as the picture — dialogue, room tone, foley, music, with their timing relationships already established.

This is the second half of the same design argument. In a chain architecture where sound is bolted on downstream, editing a line means regenerating picture and then re-syncing audio to it, which reintroduces the seam at the exact point you were trying to remove it. When the model owns both, a line change is one operation.

It also explains why the model holds up on hard material — fast rap delivery, dense syllables, breath placement mid-bar — where a two-frame drift is plainly visible and post-hoc alignment tends to fail.

How references replace prompting

The other thing that dies with the first phase is prompt folklore.

H3 takes text, images, video, and audio into one shared context. Stills fix a character’s face and a product’s exact form. A reference clip fixes movement, performance rhythm, and camera language. An audio sample fixes voice. An existing edit can carry cutting rhythm and grade forward.

So an instruction becomes comparative rather than descriptive: here’s the footage, here’s what should change, here’s the source of truth for what it should become. “The same red as before” is a hope; a reference image is a specification. Precise editing needs precise input, and that’s what the multimodal side is for.

The envelope

Clips run 5–15 seconds at 24 FPS, up to 1440p, across 21:9 through 9:16, with prompts up to 7,000 characters and reference sets of up to twelve mixed files.

Worth noting that 1440p is the level at which typography, packaging detail, and interface elements stay legible through motion — which is why H3 handles game UI, app flows, and product screens that most video models turn into letter-shaped noise the moment the camera moves.

Where the parallel breaks

Image editing tools converged on frame-exact, region-exact control. Video hasn’t got there.

Fifteen seconds is a shot, not a sequence. Preservation is strong for contained changes and degrades as the change grows — replacing a central subject or swapping an entire environment disturbs far more than removing a sign from a wall, and past a threshold you’re regenerating with extra steps. There’s no alpha, no mattes, no scene data for compositing, and no engine export. You’re directing a model, not setting keyframes.

So this is the beginning of the second phase, not its mature form. Which is roughly where image editing was when it first became genuinely useful.

Worth testing on your own footage

The evaluation that tells you anything isn’t the first render. It’s taking a clip you already have and asking for one small change.

Per-second cost sits well below comparable models, Seedance 2.0 included, which matters because iteration is how anyone converges on something good — the number of attempts you can afford sets the ceiling. Anyone assessing Minimax H3 AI Video should spend that budget on revisions rather than on fresh generations, since that’s the half of the arc the industry is now entering, and the half that decides whether a tool survives past the first week.