ChainThink reports that on July 31, according to official announcements, MiniMax released the H3 multimodal generative model, which can simultaneously understand text, images, video, and audio, and generate or edit videos based on a single natural language instruction.
H3 supports combining action migration, character reference, audio reference, and video editing into a single task. For example, users can have the model reference the camera movement from one video, a character from another image, and audio from a third clip.
This model can generate videos up to 15 seconds long, supporting 2K resolution and native stereo audio. The official API pricing is $0.13 per second for 2K, resulting in approximately $1.95 for a 15-second video;
768P costs $0.09 per second. A single task can accept up to 9 images, 3 videos, and 3 audio files, with a maximum of 12 files total.
MiniMax plans to release its model weights in the coming days; the official team has not yet announced unified evaluations with models such as Seedance and Veo.
