FEATUREDAI HOT (Curated Pool)· aihot-apiZH16:52 · 08·13
→MiniMax releases Music 3.0: open-weights music model that generates full 5-minute songs in one pass
MiniMax today released Music 3.0, an open-weights music generation model. Given a creative concept and optional lyrics, it outputs a complete song with arrangement, performance, and vocals in one pass, up to 5 minutes long. The upgrade targets three pain points: accurately interpreting creative intent, maintaining that intent across a full song, and making vocals and instruments sound performed rather than synthesized. The new Hybrid-LM architecture uses an 8B global model for song-level structure and a 0.6B local model for per-frame acoustic detail, with flow matching and a Flow-VAE converting discrete predictions to continuous audio. The post does not disclose training data scale, inference latency, the specific open-source license, or quantitative benchmark comparisons.
#MiniMax#Qwen3.5-8B
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
MiniMax open-weights a music model with an 8B+0.6B dual-model setup for up to 5-minute songs, but the post omits training data scale, latency, and the exact license.
sharp
I clicked because "open-weights" and "production-ready" rarely appear together in music generation. MiniMax Music 3.0 uses an 8B global model for song-level structure and a 0.6B local model for per-frame acoustic detail, with flow matching and a Flow-VAE converting discrete predictions to continuous audio. The idea is to apply long-sequence language modeling to music, targeting the classic problem where the first 30 seconds sound great and then the song falls apart.
I'd discount this a bit for now. The post is a product launch blog, not a technical report. No training data scale, no inference latency, no quantitative comparisons against Suno or Udio, and no clarity on what "open-weights" actually means license-wise. If it's weights with a restrictive commercial clause, that's not the same as open-source. Also unclear whether the 8B+0.6B combo runs on consumer GPUs or how long a 5-minute generation takes.
If those numbers come out and look solid, the real value is giving tooling teams a deployable, fine-tunable base model instead of being locked into API pricing. For now, treat it as a product teaser.
HKR breakdown
hook ✓knowledge ✓resonance ✓