AI MusicOpen-Source ModelsFilm ScoringHeartMuLaACE-StepYuESelf-Hosted

Creating Music with Open-Source AI Models: The 2026 Guide for Video and Film

HeartMuLa matches Suno and ACE-Step renders a 4-minute track in 20 seconds on an A100. A practical, self-hosted guide to scoring your own videos and films with open-weights models — licenses, VRAM, and a prompt-to-final-mix workflow.

September 5, 2026Michel Laclé12 min read

Why Open-Source Music Generation in 2026

You're producing videos and movies and you need background music. The default path is a subscription to a closed service — Suno, Udio, Mubert, Beatoven. Those give you tracks, but they give you a rental: a per-seat or per-track bill, an API you can't extend, models you can't fine-tune, and — critically for a creator — output you're allowed to use only on the terms the company writes and can change.

Open-weights models invert that. You download the model, you run it on your own GPU, and the output is yours under a license you can actually read. In 2026 this stopped being a novelty. Three independent releases — HeartMuLa, ACE-Step, and YuE — crossed the line from "impressive research demo" to "I would score a short film with this." One of them (HeartMuLa) now claims Suno-comparable musicality, fidelity, and controllability, and the first two are released under Apache 2.0 — the same permissive license as Kubernetes.

The short version: you can now generate original, royalty-safe background music entirely on your own hardware. If you want a permissive license and a track in 20 seconds, start with ACE-Step. If you want the highest ceiling — vocals, long-form structure, the closest open thing to Suno — go to HeartMuLa or YuE. If you need short clips, SFX, and loops rather than songs, Stable Audio Open is the clean-licensing pick.

The Landscape at a Glance

Here's the field as of September 2026, filtered to models that are genuinely open (downloadable weights) and current. Closed leaders (Suno, Udio) are excluded on purpose — they're the quality benchmark, but they're not self-hostable.

ModelTypeLicenseOutputVRAM (practical)
ACE-Step v1.5Diffusion + linear transformerApache 2.0Vocals + instrumental, to ~4 min~8 GB (turbo), 12–20 GB (XL)
HeartMuLa (oss-3B)LLM music foundation modelApache 2.0Songs w/ lyrics, section control16 GB (3B), 24 GB+ (7B)
YuE (s1-7B)LLaMA2-based full-song modelApache 2.0 + attributionFull songs w/ vocals, ~5 min16 GB min, 24 GB+ comfortable
Stable Audio OpenLatent diffusionCC-BY-NC under $1M rev.Clips / SFX, to 47 s~12 GB
Meta MusicGenAR audio language modelCC-BY-NC 4.0 (non-commercial)Instrumental, ~30 s~8 GB
Magenta RealTime (MRT2)Real-time neural synthesisApache 2.0 (code) / CC-BY (weights)Live generative instrumentalsOn-device / low

The through-line: the fast, permissive, all-round pick is ACE-Step; the quality ceiling is HeartMuLa and YuE; the non-commercial models (MusicGen, Jukebox) are for research and experimentation, not for music you ship. That last point matters more than people expect — several popular "open" music models have weights you are not allowed to use commercially.

The Models That Matter

START HERE

ACE-Step v1.5 — the permissive all-rounder

ACE-Step (from StepFun, with partners) is built to be the "Stable Diffusion moment for music" — a fast, general, flexible foundation model you can train sub-tasks on top of. Its architecture fuses diffusion-based generation with Sana's Deep Compression AutoEncoder and a lightweight linear transformer, plus semantic alignment (REPA via MERT and m-hubert) for fast convergence.

The numbers are what make it the default: it synthesizes up to 4 minutes of music in ~20 seconds on an A100 — roughly 15× faster than the LLM-based baselines — while keeping coherence and lyric alignment strong across melody, harmony, and rhythm. It handles both vocals and instrumentals, and the fine-grained acoustic detail enables voice cloning, lyric editing, and remixing.

For your use case — background music for video — ACE-Step is the highest ratio of "quality per minute of setup." Apache 2.0 means the output is commercially yours. The HF weights and the repo have a GUI, an API, and LoRA fine-tuning already wired in.

QUALITY CEILING

HeartMuLa — Suno-comparable, Apache 2.0

HeartMuLa (2026, arXiv 2601.10547) is a family of open music foundation models, not a single generator. The four components: HeartCLAP (audio–text alignment for retrieval), HeartTranscriptor (a Whisper-tuned lyrics recognizer), HeartCodec (a 12.5 Hz, high-fidelity music codec tokenizer), and the headline HeartMuLa — an LLM-based song generator that conditions on style text, lyrics, and (coming) reference audio.

Two capabilities land directly on your problem. First, fine-grained musical attribute control: you can specify the style of different sections (intro, verse, chorus) in natural language — that's how you shape a score to hit a scene's beats. Second, a dedicated "short, engaging music generation" mode the paper explicitly calls out as suitable for background music. The team reports the internal 7B version matches Suno on musicality, fidelity, and controllability, and the 3B open version is the current best open model for lyric controllability.

Caveat: it's the heavyweight. The 3B wants ~16 GB VRAM, the (unreleased) 7B ~24 GB+, and inference is currently around real-time. It's the model to reach for when a track is a hero moment, not a filler bed.

OPEN SUNO

YuE — the open full-song generator

YuE (HKUST / M-A-P) is the long-standing answer to "I want a full song from lyrics, open." It's a family of LLaMA2-based models for the lyrics-to-song task, producing structured full songs with vocals up to ~5 minutes. In the community it's the reference "open Suno."

The licensing was a real unlock: the weights are now Apache 2.0, and the team explicitly encourages artists and content creators to sample outputs into their own work, including commercial projects — the only ask is attribution ("YuE by HKUST/M-A-P") and, as always, that you own the originality and don't plagiarize. The trade-off is hardware: 16 GB is the floor, 24 GB+ comfortable, and community-quantized variants are the way to get the full 5-minute range onto consumer cards.

CLIPS / SFX

Stable Audio Open — clean-licensing short audio

Stable Audio Open (Stability AI) is text-to-audio for short material: variable-length stereo clips up to 47 s at 44.1 kHz. It's not a song model — it's a sound-design model, and that's exactly the gap it fills: ambient beds, stingers, transitions, foley, and loops you'd otherwise dig through a stock library for. Trained on a rights-respecting dataset, the license is free for commercial use under $1M annual revenue (above that you take Stability's commercial terms). If your score needs texture and SFX around a core theme, this is the cheap, clean way to source it.

RESEARCH / NON-COMMERCIAL

MusicGen, Magenta RealTime, Jukebox, Riffusion

Meta's MusicGen (AudioCraft) is mature, easy, and great for instrumental generation and fine-tuning experiments — but the released weights are CC-BY-NC 4.0, i.e. non-commercial. Don't ship a film on it. Google's Magenta RealTime (MRT2) is the pick for interactive and live generative instruments (Apache 2.0 code, CC-BY weights) — think generative ambient that responds to the scene in real time rather than a pre-rendered track. Jukebox (OpenAI, archived) and Riffusion v1 are historical/learning material now, not production tools. Note also that Google's MusicLM was never released as open weights — only its MusicCaps dataset — so it's not on this list.

Picking a Model for Film & Video Background Music

For scored video, the decision isn't really "which model is best" — it's "which failure mode can't I live with." Map your footage to a model:

  • Background beds and continuous score across many shots. ACE-Step. Fast iteration is the whole game here — you'll re-roll a bed ten times to fit a cut, and 20-second renders on an A100 (or a few minutes on a 4090) make that loop viable. Apache 2.0 keeps it royalty-free.
  • A hero theme or a titled song with real vocals. HeartMuLa or YuE. Use HeartMuLa's section-level control (verse / chorus / bridge prompts) to make the song structure follow the edit. This is your "main title" and "end credits" material.
  • SFX, stingers, and ambient texture. Stable Audio Open. Generate a 20-second impact, a riser, a room tone. Cheap, clean, and you can lay them under the core theme in your NLE.
  • Generative, reactive ambient (a scene that "breathes"). Magenta RealTime — a live instrument that stays in a key and intensity you steer. Harder to produce, but it doesn't exist in any stock library.

The practical default stack

Run ACE-Step as your workhorse for beds and general score, keep HeartMuLa (or YuE) warm for the one or two songs that carry the piece, and pull Stable Audio Open clips for SFX and transitions. Three models, one GPU, and everything you output is yours to use commercially. That's a complete, self-hosted film-music department.

Hardware: What You Actually Need

Requirements span a wide range, and the "right" answer is the model, not a fixed spec. Realistic numbers from the model repos and guides:

ModelMinimumComfortableNotes
ACE-Step v1.5~8 GB12–20 GBTurbo + small LM on 8 GB; XL/high-quality SFT on 12–20 GB
HeartMuLa oss-3B~16 GB24 GBSplit LLM/codec across 2 GPUs or use --lazy_load on one
YuE s1-7B16 GB24 GB+Community GGUF quants lower the full-song floor
Stable Audio Open~12 GB~16 GBShort clips are cheap; batch for SFX
MusicGen (3B)~8 GB~12 GBNon-commercial only

A single 16 GB card (RTX 4090-class) covers ACE-Step and Stable Audio Open comfortably and HeartMuLa-3B with lazy_load. A 24 GB card (or two GPUs, splitting LLM and codec) gets you the full stack. If you don't have that much VRAM, the honest options are: quantized/Community forks of YuE and HeartMuLa, or a hosted runner of an open model (e.g. a free web/iOS front-end that runs ACE-Step for you) — you trade pipeline control for zero setup, which is sometimes the right call for a single project.

Rule of thumb: 8 GB runs the workhorse. 16 GB runs the whole practical stack. 24 GB+ runs the quality ceiling at full length. You do not need an A100 to score a short film — you need to pick the model that fits your card.

Licensing & Commercial Use — Read This Before You Ship

This is where "open" quietly stops meaning "free to use." Model weights licenses and output permissions are different things, and they're not all the same. For a creator who will publish this music, the table below is the actual decision matrix:

ModelWeights LicenseCan I use the output commercially?
ACE-StepApache 2.0Yes. Unrestricted.
HeartMuLaApache 2.0Yes. Unrestricted.
YuEApache 2.0Yes, with attribution + originality on you.
Stable Audio OpenCC-BY-NC-styleYes under $1M annual revenue; commercial terms above.
MusicGenCC-BY-NC 4.0No. Non-commercial weights.
Magenta RealTimeApache 2.0 (code) / CC-BY (weights)Yes, with attribution.
Jukebox / RiffusionNon-commercial / OpenRAIL-MJukebox no; Riffusion v1 yes (OpenRAIL-M).

Three things to internalize. (1) Confirm the current license before shipping — these change (YuE and HeartMuLa both moved to Apache 2.0 after earlier, tighter terms). (2) "Open weights" ≠ "commercial output" — MusicGen is the classic trap: hugely popular, fully downloadable, and non-commercial. (3) You own the originality risk. Even permissive-model licenses put plagiarism/originality on you; the model reproduces statistical patterns from training data, so a track that's unmistakably one specific artist's signature is a problem regardless of license. For background beds and original themes you prompt from scratch, that risk is low — which is exactly why the prompt-driven, self-hosted approach suits a filmmaker.

A Working Scoring Workflow (Prompt to Final Mix)

Here's the loop that actually produces a usable score, not just a pile of clips. The model is the engine; the discipline is in the prompting and the DAW.

  1. Write the scene brief, not the song. Describe the function of the music: "tense, sparse, low strings, slow build, no melody, 90 seconds." For a hero theme, write actual lyrics (HeartMuLa/YuE) and tag sections [Intro] / [Verse] / [Chorus] / [Bridge] — the models parse that structure.
  2. Prompt in the model's native tag language. HeartMuLa takes comma-separated style tags (e.g. piano,tense,minor,suspense,cinematic); ACE-Step and the others take free-form style descriptions. Keep a per-scene tag string so re-rolls are reproducible.
  3. Over-generate, then select. You will not land it on prompt one. Generate 5–10 candidates per scene, listen on the actual edit (not in isolation), and keep the one that fits the cut. This is the same "render many, pick one" discipline as video generation — and it's why local speed matters.
  4. Force the instrumental where you want it. For beds, drop "instrumental, no vocals" in the style field (and leave lyrics blank / use an [Instrumental] tag) — same trick that works in the closed tools, and it works in the open ones.
  5. Layer in the NLE. Lay the core theme, then fill with Stable Audio Open SFX (risers, impacts, room tone). Duck the music under dialogue. Trim to hits. The model gives you raw material; the edit is where it becomes a score.
  6. Version and document the prompts. Save every prompt + seed that produces a keeper. When the director says "the theme but sadder," you regenerate the same seed with one tag changed instead of starting from zero.

The mental model: treat the open model like a very fast, very patient session musician. You direct the scene, it plays variations on demand, and you cut the take that fits. The entire "recording studio" now lives on your GPU and costs nothing per take.

Get Started (Minimal Commands)

The quickest path to a first track. Each is a clone, a weight download, and one generation command.

ACE-Step (fastest first result)

git clone https://github.com/ace-step/ACE-Step
cd ACE-Step
pip install -e .
# GUI + API out of the box; weights from Hugging Face (ACE-Step-v1-3.5B / v1.5)
# Then prompt: style + (lyrics or "instrumental") → ~20 s render on A100

HeartMuLa (quality ceiling)

git clone https://github.com/HeartMuLa/heartlib
cd heartlib
pip install -e .
hf download --local-dir ./ckpt HeartMuLa/HeartMuLaGen
hf download --local-dir ./ckpt/HeartMuLa-oss-3B HeartMuLa/HeartMuLa-oss-3B-happy-new-year
hf download --local-dir ./ckpt/HeartCodec-oss HeartMuLa/HeartCodec-oss-20260123
# Generate: --lyrics my_lyrics.txt --tags "piano,tense,minor,cinematic"
# Single-GPU OOM? Use --lazy_load true (or split --mula_device / --codec_device)

YuE (open full-song, vocals)

# Weights: m-a-p/YuE-s1-7B-anneal-en-icl (Hugging Face)
# Repo: github.com/multimodal-art-projection/YuE
# Apache 2.0 + attribution "YuE by HKUST/M-A-P" for public/commercial use

For any of these, start on the smallest/cheapest variant that fits your VRAM, confirm the output is usable, then scale up. If you'd rather not run a GPU for a one-off project, a hosted front-end that runs ACE-Step for you gets you finished songs with zero setup — the trade is you don't own the pipeline.

References

Written and researched by Michel Laclé. Models and licenses are as of September 2026 — reconfirm terms before shipping a project.