MiniMax launches the Music3 model to generate songs up to 5 minutes long from lyrics and descriptions.

iconMetaEra
Share
AI summary iconSummary
MiniMax has launched the Music3 model, capable of generating full songs up to 5 minutes long from lyrics and descriptions. The model employs a hierarchical autoregressive architecture, combining an 8B-parameter global language model with a 0.6B-parameter local model for acoustic details. Output is a 32kHz, 16-bit stereo WAV file that preserves musical structure and identity across all sections. The model is available on GitHub and Hugging Face with sample outputs. Crypto news continues to highlight advancements in AI and global crypto policy developments.
MiniMax has released the music generation model MiniMax-Music3, which employs a hierarchical autoregressive architecture featuring an 8B-parameter global language model responsible for modeling long-range semantics and structural changes, and a 0.6B-parameter local language model tasked with predicting acoustic codebook details. Users can input lyrics and musical descriptions to generate complete songs up to 5 minutes long, output as 32kHz, 16-bit stereo WAV files. The model preserves musical themes, rhythm, vocal identity, and arrangement progression, fully covering structures such as intro, verse, pre-chorus, chorus, bridge, instrumental interlude, and outro. The official model page and generation examples are available on GitHub and Hugging Face. This technology will further advance the field of AI music generation.

Article author and source: Hugging Face

MiniMax today launched the music generation model MiniMax-Music3, which can generate complete songs up to 5 minutes long by inputting lyrics and musical descriptions, outputting as 32kHz, 16-bit stereo WAV files.

This capability is powered by a hierarchical autoregressive architecture, with two models working in tandem, each handling distinct roles. The Global LLM, with 8 billion parameters, predicts the first RVQ codebook frame by frame, specifically modeling the long-range semantics and structural changes of songs—it is initialized from Qwen3-8B, with its embedding and output layers first adapted to musical semantic tokens, then jointly trained with the local model to model all codebooks. The Local LLM, with only 0.6 billion parameters, predicts the remaining acoustic codebooks for each frame, gradually recovering fine-grained acoustic details. One large model manages the structural framework, while a smaller model fills in the acoustic details—clearly divided responsibilities.

1.png

The official team states that the model reliably preserves the musical theme, rhythm, vocal identity, and arrangement progression in long audio tracks, accurately maintaining all structural elements such as the intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro. The official model page has been released on GitHub, along with sample generated songs for listening.

Address: https://huggingface.co/MiniMaxAI/MiniMax-Music3

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.