Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Run MiniMax H3 in ComfyUI with open weights: text-to-video, image-to-video, and reference-to-video workflows with native stereo audio, up to 2K at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit

Seedance2.0
The Future of AI Video Is Here.

Free AI Image
Truly Free AI Image Generator

Happy Horse 1.1

Gemini Omni
Gemini Omni Video Generator

Seedance 2.1
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
Why Use the comfyui minimax h3 Workflow
The comfyui minimax h3 workflow runs MiniMax's general-purpose, omni-modal generation model as open weights in ComfyUI. It jointly understands text, images, video, and audio in a single context, generating video with native stereo audio — voice, sound effects, and music modeled in one forward pass. Output reaches up to 2K resolution at 24fps for about 15 seconds, with full node-level control over every parameter.
- Native Stereo AudioDialogue, sound effects, and music generate together with the video in one MP4 — synced in a single pass through the comfyui minimax h3 workflow.
- Open-Weight ControlRun the comfyui minimax h3 model locally with full control over resolution, duration, and every diffusion parameter — no API limits.
- Multimodal Reference InputsCombine text, images, video, and audio references in one generation, locking in a character, style, motion, camera move, or voice with the comfyui minimax h3 nodes.
How to Use the comfyui minimax h3 Workflow
Start generating open-weight video with native audio in three steps using the comfyui minimax h3 workflow.
comfyui minimax h3 Workflow Features
Three native ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration — the comfyui minimax h3 workflow delivers a complete local video production stack.
Three Native Workflow Templates
The comfyui minimax h3 template library ships with text-to-video, image-to-video, and reference-to-video examples, each covering one generation mode out of the box.
Omni-Modal Context
The comfyui minimax h3 model understands text, images, video, and audio together in a single context, combining all reference types in one generation.
Reference-Driven Generation
Lock a character's identity, a style, a motion, a camera move, or a voice from reference materials — up to 9 images, 3 videos, and 3 audio clips via the comfyui minimax h3 R2V node.
Accurate Text & Brand Rendering
Spelled-out text and brand elements render cleanly with the comfyui minimax h3 model, with instruction following that describes reference relationships in natural language.
Sage Attention Speedup
Roughly double generation speed with minimal quality loss by adding the Patch Sage Attention KJ node to the comfyui minimax h3 workflow.
Resolution & Duration Grid
The comfyui minimax h3 Resolution Selector computes width and height from aspect ratio and megapixels, snapped to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.
comfyui minimax h3 — FAQ
Common questions and answers about running the MiniMax H3 model inside ComfyUI.
What is the comfyui minimax h3 workflow?
It is ComfyUI's native integration of MiniMax H3, MiniMax's general-purpose omni-modal generation model released as open weights. The workflow generates video with native stereo audio from text, images, video, and audio references in a single forward pass.
What output quality does it support?
The comfyui minimax h3 workflow outputs up to 2K resolution at 24fps for about 15 seconds. Its native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.
Which generation modes are included?
The comfyui minimax h3 template library ships with three examples: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks in character, style, motion, camera, or voice.
Does it generate audio?
Yes — the comfyui minimax h3 model produces native stereo audio including voice, sound effects, and music, modeled together with the video in one pass and synced in a single MP4 file.
How do I get started?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the pop-up to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Can I speed up generation?
Yes — install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.
Start Creating with the comfyui minimax h3 Workflow
Run MiniMax H3 locally in ComfyUI with native stereo audio, open weights, and full parameter control — text-to-video, image-to-video, and reference-to-video workflows ready to go.
