pottokao/MiniMax-H3-FL2VA-turbo-4step-v1.2-768p-NVFP4-rotated-T1

🤗 Hugging Face 来源image-to-videoapache-2.013 GBother✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo pottokao/MiniMax-H3-FL2VA-turbo-4step-v1.2-768p-NVFP4-rotated-T1 ./model-folder
需要做种者 →

MiniMax-H3 FL2VA — Turbo 4-step v1.2 (768p) · rotated-NVFP4 (T1, self-contained)

Single-file ComfyUI diffusion model for MiniMax-H3 FL2VA (First-&-Last-frame → Video + Audio), with the Turbo 4-step v1.2 768p LoRA fused in — a 4-step distilled build.

  • Self-contained: one .safetensors (base transformer + turbo LoRA fused bf16-round-first + official STRIPPED shell with curve-basis AdaLN rank-8). No separate base/LoRA needed.
  • Quantization: rotated-NVFP4, tier T1 (all-fp4 blocks), group-16 e4m3 scales, nunchaku W4A4 pack.
  • Rotation: block-256 Sylvester Walsh–Hadamard × diag(±1), seed = 100000 + blk·17 + site_id.
  • Size: ~12.5 GB — fits a single 16 GB GPU.

⚠️ Requires the custom loader

Vanilla ComfyUI cannot load this — the rotation + W4A4 pack must be decoded by:

Companions

  • Text encoder (recommended: Q3): pottokao/MiniMax-H3-TextEncoder-Qwen3VL-32B-abliterated-GGUF — use the Q3_K_S GGUF. For ComfyUI the visual-merged variant (*_vis.gguf) is required (the visual tower is what ComfyUI uses to detect the H3 text encoder).
  • MiniMax-H3 video VAE (fp16) + audio VAE (fp32)

Usage

Drop the .safetensors into ComfyUI/models/diffusion_models/, load it through the RotNVFP4 loader node, wire in the TE + VAEs. 4 sampling steps.