< Back to all clusters
[TECHNOLOGY] · Australia, Japan · 6 sources

started · updated

MiniMax launches H3 multimodal video‑generation AI model

MiniMax, an AI research company, unveiled H3, a multimodal general‑purpose model that can generate and edit short videos up to 15 seconds long in 768p or 2K resolution with native stereo audio. The system accepts text, images, video clips and audio files in a single prompt, allowing users to combine camera movements, characters, voices and music in one request. H3 can create videos from scratch, animate an initial image, define first and last frames, or transfer motion between sequences. Prompts may contain up to 7,000 characters and reference up to nine images, three video clips and three audio files, with a total of twelve resources. The model is designed for advertising, e‑commerce, branding, product design, digital interfaces, video games, film titles and animated posters. MiniMax plans to release the model as a free open‑source offering in the near future. Independent testing ranked H3 second in a text‑to‑audio‑video benchmark, behind Gemini Omni Flash and ahead of other competitors, highlighting its high quality.

The release positions MiniMax as a contender in the emerging market for integrated video‑generation AI, promising broader creative workflows for businesses and developers worldwide.

Entities

H3 · MiniMax