How to Use MiniMax H3-1
Guides & Insights

How to Use MiniMax H3: Prompts, Local Runs, and Audio That Isn't Garbage

Author

Rowan Sterling

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On 4 August, Simon Willison downloaded roughly 115 GB of weights, pointed an unofficial MLX port of MiniMax H3 at his M5 Max MacBook Pro, and asked for a rainbow-coloured skunk leaping over a mossy log in a supermarket. Just under 45 minutes later he had a clip he described as genuinely impressive — with an audio track he described as weird, speech-like garbage. His own diagnosis was the useful part: he had not given the model any audio direction, and had not read the prompting guide first. That is the shortest possible summary of what it is like to start using H3, also sold as Hailuo 3.0. The picture arrives more or less for free. Everything else — the sound, the shot timing, the 2K, the character who has to survive four shots in a row — you have to explicitly ask for, in a format that is stricter than almost any other video model in use right now.

This is a usage guide, not launch coverage. Three tiers of sourcing run through it, and they are labelled every time: the company's own documentation (the model card, the two prompt-writing guides shipped in the repository, the platform API docs); community findings from people who have actually run it — ComfyUI's maintainers, independent benchmarkers, reseller docs and Hugging Face discussion threads, all dated between 31 July and 5 August 2026; and a small number of places where we read a primary file ourselves and say so. Community numbers are single reports on uncontrolled hardware unless stated otherwise. Where practitioners disagree with each other, the disagreement is reported rather than resolved.

Before anything else: you are probably not running the whole model

Most of the confusion in the first week of H3 traces back to one structural fact that the marketing copy blurs. H3 is not a single model. According to its model card, it is a three-part system, and only the middle part was open-sourced:

H3-Context-IR — a multimodal instruction preprocessor. It reads your text, images, audio and video, reasons about how they relate, and emits a structured, semantically enriched representation of what you asked for. Hosted API only. It is exposed as its own task type that returns an enriched prompt and no video at all.

H3-Base — the 33B-parameter generation model that actually makes frames and sound. This is the open-weights part.

H3-Regenerate-2K — an in-context regeneration pass that takes the 768p result up to 2K. Hosted API only.

Two consequences follow immediately, and they decide your whole workflow. First, local generation is a 768p ceiling. H3's native canvas is a 768-pixel short edge, capped at 768×1344 — roughly 1344×768 for 16:9. If your ComfyUI output isn't 2K, nothing is broken; the package you downloaded is H3-Base, and the upscaling module is not in it. Second, the hosted API will fix a sloppy prompt and your local install will not. Everything Context-IR does for a paying API call — inferring structure, resolving which reference means what, filling in the parts you left vague — is a step you now perform by hand, or with another model, before H3-Base ever sees your words. That single asymmetry explains most reports of "the same prompt works on Hailuo but looks broken in ComfyUI."

How to Use MiniMax H3-2

Worth noting from the repository itself: the weights carry a minimax-h3-community-license-agreement tag rather than Apache or MIT (more on that below, because it is the part that decides whether you can ship at all), the model is listed at 33B parameters in F32/BF16, and as of today exactly one inference provider — fal — is listed as serving it. Twenty-three finetunes, twenty quantizations and thirty-one Spaces already exist on top of it, which is a fair signal of how fast the community moved.

The prompt format: shot blocks, not timecodes

Here is where the community and the documentation are openly in conflict, and it matters more than any other single thing in this article.

The template that spread fastest — through Reddit, then into several published guides — slices the clip into timecode brackets: [0s-2s], [2s-5s], and so on, followed by a six-block structure of style contract, timeline, camera, audio, spelled-out text and a negative list. It was reverse-engineered by reading the 45 example prompts the company shipped, and it is not nonsense: those examples really are shot lists with a sound cue sheet attached, with a median length around 130 Chinese characters and the longest running to 657 and 858 characters.

But the company's own VIDEO_PROMPT_WRITING_GUIDE_base_en.md, sitting in the repository's docs folder, prescribes something different. It organises by shot block, not by time range. We read it directly; the structure it asks for is:

Instruction — image-alignment rules, used for the keyframe modes (I2VA, FL2VA, L2VA) and omitted for pure text-to-video.

integrated_multimodal_description — the main body, divided into [Shot 1], [Shot 2], and so on.

overall_soundscape — one to four sentences of diegetic sound.

non_diegetic_music — one to three sentences of score.

Inside that, the conventions are specific enough to be checkable. [Shot 1] takes no timestamp at all — subsequent shots open with an absolute cut, phrased like "At 00:03.500, the camera cuts to…". Camera movement is written as motion type plus amplitude plus speed, folded into natural English rather than stacked as labels: a small-amplitude slow push-in, not camera: dolly-in, slow. The available motions are the familiar set — zoom, pan, tilt, tracking, arc, POV, and shake in slight or strong variants. Dialogue is wrapped in tags: <d>[English] Get in the car.</d>, with punctuation preserved exactly and never translated or paraphrased. Speakers get stable IDs — (S1), (S2), and (S1,S2) for a joint line — with age, gender, timbre and accent described on first appearance. Voiceover is marked as an off-screen line with an explicit note that the lips stay closed, which is how you stop the model from lip-syncing narration onto a face. Cross-cut conversations use a <scenetrans> marker plus a continuity statement.

The guide's explicit prohibitions are as informative as its rules: do not timestamp the first shot, do not repeat dialogue or singing or on-screen music inside overall_soundscape, do not use abstract mood words in non_diegetic_music, and never rewrite a line of dialogue. And one notable absence — the official guide has no negative-prompt section at all, which means the "banned transitions" list in the popular community template is a community invention, not a documented feature.

How seriously to take the conflict? In a Hugging Face discussion on the H3 repository, a community member posted the timecode-style structure as a guide and another practitioner replied bluntly that they had used a similar structure, that it was all wrong, and that people should read the shipped manual instead. That is one person contradicting another, not a maintainer ruling. Our reading of the evidence: both work through the hosted API, because Context-IR normalises whatever you send it; only the documented format is reliable against H3-Base directly. If you are running local weights, follow the file in the repo. If you are on the API and a timecode prompt is getting you good clips, the preprocessor is doing that work for you, and you should not conclude the format is what earned it.

Compiling a lazy brief into H3's dialect

Take Willison's twelve-word prompt as the input — a rainbow-coloured skunk leaping over a mossy log in a supermarket. Written the documented way, that becomes roughly: a [Shot 1] block that opens by naming the style (live-action, handheld, fluorescent-lit) and the frame (a supermarket aisle, moss-covered log across the linoleum, shelves receding), then the subject and the action in physical order — two padding steps, a crouch, the leap, the landing — with the camera as a slow low-amplitude tracking move that ends on the far side of the log; an overall_soundscape of claws on linoleum, the log's damp scuff, a refrigeration hum and distant trolley wheels; and a non_diegetic_music line of a short, light, plucked-string cue with no vocals. Nothing in there is creative genius. It is the same idea, with the four things H3 asks for actually supplied — and it is the difference between an unrequested room tone and a soundtrack you designed.

This step is mechanical, repetitive and a poor use of a human, which is exactly why the company shipped a tool for it that almost no coverage has picked up. The repository contains a skills directory of nine agent skills, and the first one — h3-prompt-writing — does precisely this: it takes a request and writes a structured H3 prompt across all five generation modes, complete with the soundscape and music sections. The other eight are genre recipes (product ad, 3D animated short, papercraft explainer, music-video subtitles, co-op game intro, hand-drawn live-action hybrid, and so on) that wrap the same format in a workflow.

Running that skill needs a text model, not a video model, and this is where a router is genuinely useful rather than a plug. the company's own LLM, MiniMax M3, is a sensible choice for the job and is on OrcaRouter at $0.30 per million input tokens and $1.20 per million output — provider list price, since we pass it through at 0% markup — with a 1M-token context window and, unusually, video as an accepted input type. That last detail is what makes it fit: you can hand it the official prompt guide, your brief, and an actual reference clip in one call, and get back the [Shot N] block that describes it. Compiling a prompt costs a fraction of a cent against video generation billed by the second, so there is no reason to hand-write these. Two honest caveats: H3 video generation itself does not run on OrcaRouter — the clips come from the company's platform, fal, or your own GPU — and the compile step is a convenience, not a quality guarantee. What the router buys you here is that the text half of the pipeline sits behind one key with automatic failover, so a provider outage mid-batch doesn't strand a render queue, and swapping M3 for a different model to compare compiled prompts is a string change rather than a new contract.

How to Use MiniMax H3-3

Audio is where everyone loses the first day

The single most corroborated failure mode in the first week is the one Willison hit: leave the sound unspecified and H3 invents something, badly. He got speech-like noise from a prompt with no audio direction. The reverse-engineering write-up of the official examples reports the same class of failure in tamer form — omit the audio block and the model ships an unrequested room tone — and frames these as absences rather than mistakes, which is the right way to think about a joint audio-video model. It is always generating a soundtrack. Your only choice is whether you specify it.

Pooling the guidance in the official prompt guide with what reseller documentation and reviewers report actually working:

Dialogue is the hardest thing to ask for. Keep spoken lines to one or two sentences per five seconds of clip. Overlong lines produce rushed delivery, audio that runs past the last frame, or strained lip-sync — reported consistently across reviewers.

Describe the speaker before the line, and the delivery with it. Age, gender, timbre and accent on first appearance per the official guide; simple, direct delivery notes (clearly, warmly, flatly) rather than elaborate performance direction, which reportedly degrades sync.

Bind every effect to a visible event. "A cork pop exactly as the cork flies" rather than "sound: a pop". The words as, when and then are what carry the synchronisation.

Brief music by genre, tempo, mood and instrumentation — never by artist or song. Add "no vocals" when the picture should lead; it is a documented way to keep the mix clean.

Layer ambience in one clause. Rain on glass, a low murmur of conversation, the occasional cup clink, soft background jazz — stacked in a single sentence rather than itemised.

Do not restate dialogue in the soundscape section. An explicit don't in the official guide; the sections are meant to be disjoint, and repeating a line there is a way to get it twice.

If a word must be readable on screen, type the word in quotes. Reviewers report that specified strings render cleanly while vague requests ("HUD elements", "a sign") come back as letter-shaped noise.

Where the audio genuinely delivers, per multiple reviewers, is concrete physical sound — foley and effects tied to something visible — and ambience. It is weaker on abstract or atmospheric requests. Multi-speaker scenes frequently need editing to get speaker order right, and pronunciation quality varies by language, so test your target language on a throwaway clip before committing a project to it. Several reviewers also note the honest ceiling: the sound is good enough for social distribution and generally not good enough to ship as a broadcast or paid-ad mix without replacement. None of that is vendor guidance; it is what practitioners report.

References: give every file a job

H3's reference system is its strongest differentiator and the place where setup mistakes are most common. The documented limits, from the platform API docs: up to 9 images, 3 video clips and 3 audio clips, capped at 12 files total, with each reference video or audio between 2 and 15 seconds and the total across them not exceeding 15 seconds. Reference-to-video requires at least one image or video; audio alone is not accepted.

The technique that both the official reference guide and community reports converge on is to tag each input and assign it a job. Reference by tag in the exact order the files were attached — <Picture 1>, <Video 1>, <Audio 1> — and then state in the prompt which reference drives which property: identity, style, motion, camera, or voice. The official example prompts open exactly this way, declaring that image one is the overall mood and style reference and image two is the lead character. Practitioners report explicit assignment working substantially better than dumping nine images and hoping.

Two details that cost people renders:

<Picture 1> is not a mood board — it is the literal first frame at 0.000 seconds, and it belongs to [Shot 1]. Describe what is in it (style, subjects, composition, scene anchors) before describing the action that follows from it. Treat it as a still you are animating, not a hint.

There are two separate checkpoints, and using the wrong one silently fails. fl2va handles text-to-video and first/last-frame work; ref2va handles reference-to-video. This is the most-reported ComfyUI error: the R2V graph runs with the FL2VA model still selected. Worth knowing too that a commenter on the ComfyUI announcement points out the shipped R2V template exposes only two image reference slots even though the model accepts nine — a template limitation, not a model limit.

On the hosted side, two more practitioner notes from reseller documentation: the ratio parameter is required on ref2va endpoints and cannot be left as "adaptive", and overlapping jobs return a 429 for task concurrency rather than queueing — so batch sequentially, or build your own queue. Locally, ref_image_size defaults to match, which is faster; max preserves up to a 2048-pixel short edge on the reference and costs time.

What a local run actually costs in wall-clock time

Downloading the full official repository is 498 GB, and the ComfyUI-repackaged mirror is 343 GB. You do not need either in full. The four files ComfyUI's own tutorial lists for text-to-video and image-to-video come to about 42.5 GB:

Diffusion modelminimax_h3_fl2va_pruned_int8_convrot.safetensors, 20.97 GB, into models/diffusion_models.

Text encoderqwen3vl_32b_minimax_h3_nvfp4_awq.safetensors, 15.69 GB, into models/text_encoders.

Video VAEminimax_h3_video_vae_fp16.safetensors, 5.21 GB, into models/vae.

Audio VAEminimax_h3_audio_vae_fp32.safetensors, 0.61 GB, also into models/vae.

Add for reference-to-videominimax_h3_ref2va_pruned_int8_convrot.safetensors, another 20.97 GB.

That 42.5 GB is not a VRAM figure — it is disk, and the pieces are streamed and offloaded. The reason the model fits on consumer cards at all is an engineering trick ComfyUI documented: roughly 40% of the parameters live in AdaLN modulation branches whose outputs can be precomputed and replaced by functionally equivalent lookup tables, taking the loaded model from 33.12B to about 20.11B with, per ComfyUI, no quality loss. Full precision would be 123.6 GB. Unpruned int8 checkpoints (34.04 GB each) and bf16 ones (66.28 GB each) exist if you are fine-tuning or chasing the last few percent of fidelity.

ComfyUI 0.30.0 or later is required, and it ships three templates: T2V, I2V and R2V. The baseline every guide and maintainer converges on for a first run is deliberately small: 16:9 at 0.4 MP (864×480), 5 seconds, 20 steps, res_multistep sampler with the simple scheduler, denoise 1.0, batch 1. Get one MP4 out that contains both video and stereo audio before you touch resolution or length. One arithmetic quirk to expect: durations snap to a 17k+5 frame grid, so a 5-second request becomes 124 frames — about 5.17 seconds at 24 fps — and a 10-second one lands at 243 frames, or 10.125 seconds. Requests in the 5–15 second range are the safer zone.

How to Use MiniMax H3-4

What that costs in real time, from community reports (all single runs, uncontrolled, on different settings — the chart above prints each configuration beside its bar): a 12 GB RTX 3060 with 32 GB of system RAM and a fast NVMe completed the 864×480 / 124-frame / 20-step baseline in under nine minutes using dynamic offload. A 16 GB RTX 4090 laptop with SageAttention enabled did 960×540, 5 seconds, 20 steps in 182 seconds. An NVFP4 build on a single RTX 5090 produced a 243-frame (10.1 s) 864×480 clip at 10 steps in 175 seconds, peaking at 26.9 GiB of VRAM with just 31.7 GB of files on disk. SGLang's own cookbook reference — 1344×768, 124 frames, 50 steps across two RTX 5090s with layerwise offload — took 559.67 seconds. And the Apple Silicon path, via the unofficial MLX port, took just under 45 minutes for a single clip.

Practical notes distilled from those reports:

System RAM and disk speed are load-bearing, not optional. On a 12 GB card the selected files exceed VRAM by design, so the offload path runs through RAM and then the SSD. 32 GB of RAM appears in both successful low-VRAM reports. No verified 8 GB result exists.

SageAttention is the one free speedup — roughly 2× per ComfyUI, less if your run is dominated by offload traffic. Install the wheel matching your exact PyTorch and CUDA build plus ComfyUI-KJNodes, then insert Patch Sage Attention KJ between the UNET loader and the guider and set it to auto. Expect warnings that some H3 layers are not FP16/BF16; that fallback is documented and normal.

Steps are negotiable. The 5090 benchmark ran 10 and the ComfyUI baseline runs 20, while the SGLang reference config uses 50. Nobody has published a quality-per-step curve, so find your own floor before you assume you need 50.

Change one variable at a time out of OOM. The documented recovery is to drop back to 0.4 MP, 5 seconds, batch 1, and move a single knob per attempt.

Silent audio means the audio VAE isn't wired to the video output node. A distinct failure from bad audio, and easy to misread as the model refusing to generate sound.

Apple Silicon is unofficial. Willison's route was PipeNetwork/minimax-h3-mlx, a community port, run via uv against an MLX requirements file. the company's own materials name SGLang, vLLM, Diffusers and ComfyUI, with a typical deployment of four GPUs in BF16, and say nothing about Metal or MPS. Treat Mac support as community-maintained.

Local versus hosted, with the numbers

Because the 2K module and the prompt preprocessor are hosted-only, this is not really an either/or. The pattern the architecture pushes you toward is: iterate locally at 768p where each attempt costs electricity and three minutes, then buy the finish. A per-second charge is fine for the take you keep and ruinous for the forty you throw away.

On price, be careful, because it varies about 2× by who you buy from and the vendor's own blog post quotes no dollar figure — only that 2K comes in at less than a third of mainstream models and 768p at under half the price of rivals' 720p. What is concretely published: fal, currently the only inference provider listed on the Hugging Face repo, charges $0.16 per second at 768p and $0.26 per second at 2K. Secondary trackers report the company's own list price materially lower — around $0.09–$0.10 per second at 768p and $0.13–$0.14 per second at 2K — and they do not agree with each other to the cent, so treat those as indicative and check the platform pricing page before you model a budget. Also budget for the fact that a reference video is billed on its own duration as well as the generated output.

Run it through a real number. A hundred keeper clips at eight seconds each is 800 generated seconds: about $112 at a $0.14 2K rate, roughly $208 at fal's $0.26, and around $76 if you deliver at 768p on the low list price. The forty rejects per keeper are what decide the bill, and those are the ones to generate locally. This is also the reason to keep the price you are comparing honest — on OrcaRouter, the text side of that pipeline is passed through at provider list with 0% markup, so when a vendor cuts a price it is live the same day rather than after a margin recalculation. Video generation, again, isn't ours; the point is only that the compile step should not be the line item you have to think about.

Read the license before you build a product on this

This is the part most "how to use" guides skip, and for a lot of readers it is the only part that changes a decision. We read the LICENSE file in the repository directly. The weights ship under the H3 Community License Agreement, not an open-source licence, and it contains terms that are unusual enough to quote in substance:

Territory. The licence defines its "Applicable Territory" as worldwide excluding the European Union, the United Kingdom, the Republic of Korea and the United States of America. Read plainly, that excludes a large share of the people currently posting local-generation results — and the restriction is written to reach the outputs, not only the weights.

Attribution. Commercial use requires displaying "H3" on the product's interface; the licence also encourages a "Powered by H3" notice.

Revenue threshold. Organisations earning over $20 million a year from products or services built on it must obtain separate written authorisation from the company.

No distillation. You may not use H3 or its outputs to improve any other AI model, except H3 derivatives. That rules out the standard synthetic-data play.

Governing law. Hong Kong SAR, with exclusive jurisdiction in Hong Kong courts.

One genuinely open component. The Qwen3-VL-32B text encoder is Apache 2.0; the restrictions above attach to the H3 weights.

We are not your lawyers and this is not legal advice — read the licence and the Q&A document the company ships beside it, and if you are building commercially in an excluded territory, get counsel rather than a blog's reading. The practical fork: the hosted API is a different transaction with different terms from the community licence on the weights, so if the local licence doesn't work for you, the API route may still be available. Check the terms of whichever platform you buy from.

A first-week test plan

Five runs, in this order, is enough to know whether H3 belongs in your pipeline:

Run 1 — prove the plumbing. T2V template, 864×480, 5 seconds, 20 steps, res_multistep/simple. Success criterion is an MP4 with stereo audio in it, not a good clip.

Run 2 — prove the format matters. Same subject twice: once as a loose one-line description, once written as [Shot 1] plus overall_soundscape plus non_diegetic_music. If the second isn't clearly better against H3-Base, something in your setup is wrong.

Run 3 — one line of dialogue. One speaker, one sentence, delivery described, five seconds. This is the fastest read on whether the audio is good enough for your use, and on how your target language is pronounced.

Run 4 — reference discipline. R2V with the ref2va checkpoint, two or three references, each explicitly tagged and given a job, with <Picture 1> described as the actual first frame. Then break it deliberately: attach the same references with no assignments and compare.

Run 5 — the finish. Take your best local 768p prompt to the hosted API at 2K and see what Context-IR and the regeneration pass add. That delta is what the per-second price is buying, and it is the only way to price the hybrid workflow honestly.

FAQ

Can I get 2K out of the local weights?

No. The open package is H3-Base, whose native canvas is a 768-pixel short edge (up to 768×1344). 2K comes from H3-Regenerate-2K, a separate in-context regeneration module that the company kept on the hosted API. You can upscale locally with any general-purpose upscaler, but that is a different operation from H3's own regeneration pass and will not match it.

Why does the same prompt behave differently on the API and in ComfyUI?

Because the API runs H3-Context-IR first. It reads your text and your references, reasons about how they relate, and hands H3-Base a structured, enriched instruction. Locally there is no such stage — H3-Base gets your raw words. A vague prompt that produces a competent clip through the API is being rescued by a preprocessor you did not install, which is why the documented prompt format matters far more to local users than to API users.

Is the [0s-2s] timecode template wrong?

It is not what the shipped guide asks for. The company's VIDEO_PROMPT_WRITING_GUIDE_base_en.md organises by [Shot N] blocks, gives Shot 1 no timestamp, and expresses later cuts as absolute times ("At 00:03.500, the camera cuts to…"). It also has no negative-prompt section, so the "banned transitions" list in the popular template is a community addition. That said, practitioners are reporting good API results with the timecode form — plausibly because Context-IR normalises it. Our reading: use the documented format against local weights, and don't credit the timecode brackets for API results the preprocessor may have earned.

Can a company in the US or the EU use the open weights commercially?

The licence text we read excludes the EU, the UK, South Korea and the US from its Applicable Territory, and the exclusion is written to cover outputs as well as weights. That is a serious obstacle for a commercial deployment in those places, and it is a legal question rather than a technical one — read the licence and the Q&A file yourself and take advice. The hosted API is governed by the platform's own terms, not the community licence, so it is worth evaluating separately rather than assumed to be equally restricted.

What to check before you commit

H3 is unusually good value for a specific shape of work: short, sound-designed, reference-consistent clips where you are willing to write a real shot list. It is unusually demanding about how you ask. The three things worth verifying yourself, because they are the ones that move fastest: whether more inference providers appear beside fal (which is what will bring the per-second price down), whether the ComfyUI R2V template grows past two reference slots to the model's nine, and whether the community's timecode habit or the documented shot-block format wins once someone runs a controlled comparison against local weights. Nobody has published that comparison yet. Until they do, the manual in the repository is the better bet — and it is sitting in the same download you already spent 42 GB on.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube