joeygambino/MiniMax-H3-Multishot-Workflow

🤗 On Hugging Facetext-to-videomit3 MBother✓ Checksum-verifiedupdated 0d ago
Magnet

MiniMax-H3 Multishot Workflow

Chain multiple MiniMax-H3 shots into one continuous video with audio - in

one node. Each shot starts from the last frame of the previous one; the

duplicated seam frame and its 1/24s of audio are trimmed automatically.

The demo below was made by this workflow: 30 seconds, three chained shots

from one script, same presenter and same voice across both seams, rendered on

the Q5_1 GGUF.

Files

  • H3_Multishot_AIO.json - easy mode: loaders > script box > one sampler

node > save. Write one prompt per shot with --- between them, pick a

shot count (0 = one shot per prompt, 1-8 forces it), queue.

  • H3_Keyframes.json - keyframes at any position, not just first/last.

Up to 6 anchor images placed by fraction (0, 0.5, 1) or absolute frame

index, in a single generation - so the audio is one continuous stream.

  • H3_Multishot_MEMORY.json - long form (2-5 minutes). Keeps a persistent

identity anchor so drift cannot compound across many shots.

Requirements

  • ComfyUI v0.30.0+ (native MiniMax H3 support)

includes the one-line ComfyUI-GGUF architecture patch)

for 16 GB) or the originals; text encoder + VAEs from

Comfy-Org/MiniMax-H3

Notes

  • frames_per_shot sits on H3's 17k+5 frame grid (243 = ~10.1s at 24 fps;

362 = ~15.1s, the trained max).

  • End every shot on what the NEXT shot expects to see - the chain hands each

shot the previous final frame, and matching that bridge to the next shot's

framing is what makes seams invisible.

  • Malformed JSON scripts fail loudly instead of rendering the raw text.

v1.2 (2026-08-04)

  • Keyframes at any position. Stock ComfyUI pins H3 keyframes to the first

and last frame and raises only first/last keyframe anchors are supported

for anything else. That is a positional-maths limit, not a model limit -

both stock cases are the same expression:

  cond_t = text_len + FRAME_RESCALE * pixel_index

which is defined for every frame. Measured on an RTX 5090, 243 frames, one

anchor at pixel frame 121: the rendered frame closest to the anchor image was

frame 122 - off by one - reached by continuous motion with no cut, audio

unbroken through it. A three-anchor run landed frames 121 and 242 exactly.

Applied in memory; it does not edit any ComfyUI file, self-tests against the

stock formula, and rolls back if first/last positions do not reproduce.

  • Move vs cut. Anchor images with a plausible camera path between them make

H3 interpolate. Images with no possible path (a kitchen and a diner) make it

cut, then hold - stock first/last does the same with such a pair, so that is

the model, not the node. The cut case is still useful: a timed shot change

inside ONE generation, which keeps the audio continuous across it.

  • Two new controls: H3 Condition Strength (exposes

minimax_visual_cond_noise_aug / minimax_audio_cond_noise_aug, read by

ComfyUI core but written by no stock node) and H3 Reference Audio, a

stereo guard - a mono reference clip crashes the sampler with an unhelpful

shape mismatch because the layout reserves two channels.

v1.1 (2026-08-04)

  • Image-to-video. The Multishot Sampler now takes an optional start_image.

Connect a LoadImage and shot 1 starts from that frame, then chains as usual.

Leave it unconnected for unchanged text-to-video behaviour - v1.0 graphs keep

working.

  • ~4x faster on 32GB cards. The text encoder is evicted before sampling.

The Qwen3-VL encoder (~16.5GB even at Q4) and the H3 DiT (~25GB) do not co-fit

on 32GB, so the DiT was loading partially and streaming ~19GB from system RAM

every step. Measured on an RTX 5090: ~60 min -> ~15 min for the same render.

  • Text encoder GGUF + mmproj published:

joeygambino/MiniMax-H3-encoder-GGUF.

The mmproj sidecar is required for multi-shot, not just for reference

images - shot chaining feeds the previous frame through the encoder's vision path.

  • Render at H3's native resolution and upscale afterwards: native 1920x1088

measured worse than 960x544 in blind review and cost ~4x the time.

Node pack: ComfyUI-H3-Multishot

On Civitai

Support

Everything here is free and stays free. If it saved you a night of debugging,

tips keep the 5090 warm: