comfyui minimax h3
Produce synchronized stereo-audio clips through the comfyui minimax h3 workflow
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Create open-weight videos with synchronized stereo audio using the comfyui minimax h3 workflow. Supports text, image, and reference prompts at up to 2K/24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why the comfyui minimax h3 Pipeline Is a Game Changer

The comfyui minimax h3 pipeline brings MiniMax's omni-modal open-weight model into ComfyUI, letting you work with text, images, video, and audio in one shared context. It produces clips with naturally synced stereo audio, including voices, effects, and music generated in a single pass. Outputs scale to about 15 seconds at 2K resolution and 24fps, with fine-grained control over every parameter at the node level.

  • Built-In Stereo Sound
    Dialogue, sound effects, and music are generated alongside the image track in a single MP4, perfectly synced by the comfyui minimax h3 workflow.
  • Full Node-Level Freedom
    Run the comfyui minimax h3 engine locally and adjust resolution, duration, and any diffusion parameter directly — no external API constraints.
  • Rich Multimodal References
    Combine text prompts with images, video snippets, and audio cues in one go to lock a character, style, motion, camera angle, or voice through the comfyui minimax h3 nodes.

Getting Started with the comfyui minimax h3 Workflow

Produce open-weight videos with synced audio in three simple steps using the comfyui minimax h3 workflow.

Inside the comfyui minimax h3 Workflow Toolkit

The comfyui minimax h3 workflow pairs three built-in ComfyUI templates, open-weight multimodal generation, native stereo audio, reference-driven control, and optional Sage Attention acceleration for a complete local production pipeline.

Three Ready-Made Templates

The comfyui minimax h3 template set includes text-to-video, image-to-video, and reference-to-video examples, with each mode ready to run immediately.

Unified Omni-Modal Context

The comfyui minimax h3 engine interprets text, images, video, and audio as one connected context, so you can blend all reference types within a single generation.

Reference-Driven Creativity

Lock a character's identity, art style, motion, camera movement, or voice from inputs — up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 reference node.

Clean Text & Brand Rendering

The comfyui minimax h3 model handles spelled-out text and brand elements accurately, and it follows natural-language instructions about reference relationships faithfully.

Sage Attention Acceleration

Add the Patch Sage Attention KJ node to the comfyui minimax h3 workflow to nearly double generation speed with minimal quality trade-off.

Flexible Resolution & Duration Grid

The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapping to the model's 32-multiple canvas and 17-frame-per-block timing at 24fps.

FAQ

comfyui minimax h3 — Common Questions

Answers about running the MiniMax H3 open-weight model inside ComfyUI for video generation.

1

What exactly is the comfyui minimax h3 workflow?

It is ComfyUI's built-in support for MiniMax H3, a general-purpose omni-modal model released as open weights. The workflow produces video with native stereo audio from text, image, video, and audio references in one forward pass.

2

What resolution and frame rate can I expect?

The comfyui minimax h3 workflow reaches up to 2K at 24fps for roughly 15 seconds. Its native canvas runs at a 768px short edge, limited to 768x1344 pixels and rounded to multiples of 32.

3

Which generation modes come with the template?

The comfyui minimax h3 template library contains three modes: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that preserves character, style, motion, camera, or voice.

4

Does it really generate audio?

Yes — the comfyui minimax h3 model creates native stereo audio, covering voice, sound effects, and music. That audio is modeled together with the video in one pass and synced into a single MP4 file.

5

How do I start using it?

Update ComfyUI to 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 template, and follow the pop-up to download models from the Comfy-Org/MiniMax-H3 Hugging Face repository.

6

Is there a way to speed things up?

Absolutely — install SageAttention and KJNodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double generation speed.

Begin Producing with the comfyui minimax h3 Workflow Today

Run MiniMax H3 locally in ComfyUI and enjoy native stereo audio, open weights, and total parameter control — with T2V, I2V, and R2V templates ready to use at once.