MiniMax-H3Video GenerationLocal AIstable-diffusion.cppGGUFRTX 3090Audio GenerationCUDA

Running MiniMax-H3 Locally: Unified Audio+Video Generation

How a 134 GiB unified audio+video generation model runs cloud-free on a 4x RTX 3090 host via GGUF quants and stable-diffusion.cpp — with the Q4/Q6/Q8 sample renders you can watch.

August 30, 2026Michel Laclé7 min read

What MiniMax-H3 Is

MiniMax-H3 is a unified audio-and-video generation model: one model produces both a video stream and a synchronized audio stream in the same diffusion pass. It's not "video, then tack on sound" — the audio is generated jointly from the same latent space, which is what makes it interesting as a compact, self-contained media generator [Official weights].

The catch is size. The official release is a full-BF16, sharded deployment: a 61.7 GiB FL2VA transformer, a 62.1 GiB Qwen3-VL-32B text encoder, a 9.7 GiB video VAE, and a small audio VAE — about 134 GiB per variant. MiniMax's own SGLang example runs it on four 80 GiB A100/H100 GPUs [HF].

The whole point of this work: that 134 GiB release doesn't fit on consumer hardware. But the GGUF-quantized versions of MiniMax-H3 (via unsloth/MiniMax-H3-GGUF) do — and they can be rendered locally with stable-diffusion.cpp. This repo is the deployment + render tooling that makes that a one-command, config-first workflow.

Why Run It Locally

Four reasons this is worth running yourself rather than calling an API:

  • No per-second cloud billing. Quality iteration means rendering the same scene many times at different resolutions and step counts. Locally that's free compute; in the cloud it's a line item that adds up fast.
  • Deterministic A/B testing. The repo's discipline is "change one factor at a time" — resolution, then steps, then prompt, then conditioning. Same prompt, same seed, same model family every time, so a quality delta is real signal, not noise.
  • Full ownership of the pipeline. Every render writes a metrics.tsv phase log (asset cache, conditioning, generate_video, muxing), so you can measure and optimize the exact bottleneck.
  • A real quant ceiling to push. The 4x RTX 3090 host (96 GiB aggregate) sits below the ~160 GiB effective floor for the official BF16 release [BF16 viability note]. The GGUF quants are the only way to run it on that hardware — and finding which quants actually load is part of the work.

The Local Stack

The render path is stable-diffusion.cpp (the CUDA-enabled sd-cli), not the llama.cpp chat server. A render is a fixed pipeline:

  1. Asset availability — check/download the GGUF denoiser, text encoder, and VAEs from the Hugging Face cache.
  2. Text conditioning — Qwen3-VL text encoder on CPU/RAM + CUDA staging.
  3. AV latent sampling — the MiniMax-H3 denoiser with CUDA flash attention.
  4. Video/audio decode — the MiniMax video + audio VAEs.
  5. WebM mux — VP8 video copy + Opus audio encode into a single shareable clip.

Two conditioning modes are exercised: FL2VA (first-and-last-frame to video+audio) and Ref2VA (reference image to video+audio). The model requires 17k + 5 frames, so a requested 5-second scene at 24 FPS is aligned up to 124 frames (~5.17 s).

The Q8 compatibility gotcha

Not every GGUF in the wild loads. The Abiray full and pruned Q8 artifacts both fail stable-diffusion.cpp metadata validation — their adaln_proj / time_embedder tensor layouts don't match the runtime. The working Q8 is the unsloth minimax_h3_fl2va_pruned-Q8_0.gguf family (21.4 GB denoiser + matching 18.2 GB Q4 text encoder + VAEs). That artifact family is isolated under its own model cache so it can't be mixed with the Q4/Q6 files [Q8 compatibility note].

The Quant Ladder

The renders span a quantization and quality ladder, each a controlled step up from the last:

Test case Quant Conditions generate_video
FL2VA smokeQ4320x192, 4 steps~94 s / seg
Ref2VA smokeQ6320x192, 4 stepshistorical
FL2VA qualityQ8864x480, 20 steps, 5.17 s691.6 s
FL2VA quality rungQ81024x576, 20 steps, 5.17 s1,082.9 s
FL2VA detailed humansQ81024x576, 20 steps, 5.17 s1,087.0 s

The smoke tests (Q4/Q6, 320x192, 4 steps) exist to verify a quant loads and produces valid WebM. The Q8 rung tests are the actual quality development — resolution and step count pushed up one factor at a time, then a detailed-prompt test to isolate conditioning from sampler changes. All renders use seed 42, 24 FPS, and --cfg-scale 1.0.

Watch the Renders

These are the actual sample clips, rendered on the 4x RTX 3090 host. Start at the top (fast Q4 smoke) and work down to the 1024x576 detailed scene.

1. FL2VA — Q4 smoke test

320x192 · 4 steps · first/last-frame conditioning · ~30 s · 9.0 MB

2. FL2VA — Q4 smoke test (run 2)

320x192 · 4 steps · reproducibility check · ~30 s · 6.9 MB

3. Ref2VA — Q6 reference-to-video

320x192 · 4 steps · reference-image conditioning (Q4 text encoder) · ~30 s · 6.4 MB

4. FL2VA — Q8 smoke test

320x192 · 4 steps · verifies the 21.4 GB unsloth Q8 denoiser loads · ~30 s · 9.1 MB

5. FL2VA — Q8 quality development

864x480 · 20 Euler steps · 5.17 s · seed 42 · 6.5 MB

6. FL2VA — Q8 quality rung (higher res)

1024x576 · 20 steps · 5.17 s · seed 42 · 7.5 MB

7. FL2VA — Q8 detailed human fighters

1024x576 · 20 steps · 5.17 s · detailed prompt (same resolution/steps as #6) · 10 MB

Throughput, Honestly Stated

The numbers are slow, and that's the real constraint. The generate_video phase — the part you can't shortcut — runs at roughly:

  • ~6.3 s of render per second of video at the 320x192 / 4-step smoke conditions (fast).
  • ~134 s of render per second of video at 864x480 / 20 steps.
  • ~210 s of render per second of video at 1024x576 / 20 steps.

A single 5.17-second quality scene at 1024x576 takes about 18 minutes of generation. That's fine for a controlled quality test where you're changing one variable; it's a hard ceiling for anything that needs to iterate in real time. The measurements are the generate_video phase only — they exclude model download and full process wall time — and they're the baseline for future optimization (better quant, more VRAM, or a larger GPU).

The design philosophy

The repo does one thing at a time and keeps every parameter for a given target in a single config file (targets/<name>/<name>.env). The implemented target is dev-rtx4070 (a single 8 GB card, Q2_K — the smallest quant that fits, for dev/test only); the 4x RTX 3090 render host is where the quality samples above came from. A future production target drops in alongside them without touching the existing ones. It's the same "one repo, one model, config-first" philosophy as the 4x RTX 3090 Qwen 3.8-27B deployment [Repo].

Reproduce It

The render entrypoints (on a host with the CUDA-enabled sd-cli and the model cache populated):

# Verified Q4 render (fast smoke test)
bash create-stick-fighter-video-q4.sh

# Reference-to-video Q6
bash create-stick-fighter-video-q6.sh

# Q8 quality development (864x480, 20 steps, seed 42)
bash create-stick-fighter-video-q8.sh

# Next rung: 1024x576 at the same 20-step conditions
bash create-stick-fighter-video-q8-1024x576.sh

# Detailed human-fighter scene at the proven 1024x576 level
bash create-detailed-matrix-fight-video-q8.sh

# Or run the two verified configurations back to back
bash render-quant-tests.sh --all

Every render lands under ~/videos/minimax-h3/<quant-name>/ with a metrics.tsv next to the WebM. Model weights stay under ~/models/minimax-h3-render*/ and are git-ignored.

Bottom line: a 134 GiB unified audio+video generation model is genuinely runnable on consumer multi-GPU hardware — not the official BF16 release, but its GGUF quantized forms via stable-diffusion.cpp. You trade the A100's speed for full local ownership: free, deterministic, measurable, and repeatable. The ceiling is render time, not capability.


References

  1. michellacle/video-streaming-minimax-h3 — the deployment + render tooling
  2. MiniMaxAI/MiniMax-H3 — official BF16 release
  3. unsloth/MiniMax-H3-GGUF — the working GGUF quant family
  4. leejet/stable-diffusion.cpp — the CUDA render engine
  5. 4x RTX 3090 Qwen 3.8-27B — sibling deployment repo

Research by ThinkSmart.Life · August 2026

↑ Back to top