Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Create open-weight videos with synchronized stereo audio using the comfyui minimax h3 workflow. Supports text, image, and reference prompts at up to 2K/24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator
Why the comfyui minimax h3 Pipeline Is a Game Changer
The comfyui minimax h3 pipeline brings MiniMax's omni-modal open-weight model into ComfyUI, letting you work with text, images, video, and audio in one shared context. It produces clips with naturally synced stereo audio, including voices, effects, and music generated in a single pass. Outputs scale to about 15 seconds at 2K resolution and 24fps, with fine-grained control over every parameter at the node level.
- Built-In Stereo SoundDialogue, sound effects, and music are generated alongside the image track in a single MP4, perfectly synced by the comfyui minimax h3 workflow.
- Full Node-Level FreedomRun the comfyui minimax h3 engine locally and adjust resolution, duration, and any diffusion parameter directly — no external API constraints.
- Rich Multimodal ReferencesCombine text prompts with images, video snippets, and audio cues in one go to lock a character, style, motion, camera angle, or voice through the comfyui minimax h3 nodes.
Getting Started with the comfyui minimax h3 Workflow
Produce open-weight videos with synced audio in three simple steps using the comfyui minimax h3 workflow.
Inside the comfyui minimax h3 Workflow Toolkit
The comfyui minimax h3 workflow pairs three built-in ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration for a complete local production pipeline.
Three Ready-Made Templates
The comfyui minimax h3 template set includes text-to-video, image-to-video, and reference-to-video examples, with each mode ready to run immediately.
Unified Omni-Modal Context
The comfyui minimax h3 engine interprets text, images, video, and audio as one connected context, so you can blend all reference types within a single generation.
Reference-Driven Creativity
Lock a character's identity, art style, motion, camera movement, or voice from inputs — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 reference node.
Clean Text & Brand Rendering
The comfyui minimax h3 model handles spelled-out text and brand elements accurately, and it follows natural-language instructions about reference relationships faithfully.
Sage Attention Acceleration
Add the Patch Sage Attention KJ node to the comfyui minimax h3 workflow to nearly double generation speed with minimal quality trade-off.
Flexible Resolution & Duration Grid
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapping to the model's 32-multiple canvas and 17-frame-per-block timing at 24fps.
comfyui minimax h3 — Common Questions
Answers about running the MiniMax H3 open-weight model inside ComfyUI for video generation.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's built-in support for MiniMax H3, a general-purpose omni-modal model released as open weights. The workflow produces video with native stereo audio from text, image, video, and audio references in one forward pass.
What resolution and frame rate can I expect?
The comfyui minimax h3 workflow reaches up to 2K at 24fps for roughly 15 seconds. Its native canvas runs at a 768px short edge, limited to 768x1344 pixels and rounded to multiples of 32.
Which generation modes come with the template?
The comfyui minimax h3 template library contains three modes: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that preserves character, style, motion, camera, or voice.
Does it really generate audio?
Yes — the comfyui minimax h3 model creates native stereo audio, covering voice, sound effects, and music. That audio is modeled together with the video in one pass and synced into a single MP4 file.
How do I start using it?
Update ComfyUI to 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 template, and follow the pop-up to download models from the Comfy-Org/MiniMax-H3 Hugging Face repository.
Is there a way to speed things up?
Absolutely — install SageAttention and KJNodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.
Begin Producing with the comfyui minimax h3 Workflow Today
Run MiniMax H3 locally in ComfyUI and enjoy native stereo audio, open weights, and total parameter control — with T2V, I2V, and R2V templates ready to use at once.
