Now live

MiniMax H3: Video with Native Sound

A clip from the MiniMax H3 launch showcase. Source: MiniMax official demos.

What the model puts out

Every clip below comes from MiniMax’s own launch demos, so you can judge the picture quality before spending anything. They loop silently here — native stereo sound is covered further down.

Five things it does well

The claims are MiniMax’s own and the clips are their demos. What you get on A2E is whatever your own generation returns.

Native multimodal understanding & generation

Accepts text, images, audio, and video as inputs. Understands characters, motion, sound, emotion, camera work, style, and intent, and fuses multiple references into unified audiovisual output. MiniMax’s own example is a single sentence: “Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.” Three kinds of material, each given one job, in one prompt.

Precise multimodal editing & control

Edits characters, objects, scenes, sound, and pacing across multiple dimensions with fine-grained instruction following, so you can iterate on existing content predictably. Changing one thing does not mean regenerating everything: say what has to stay and what has to move, and carry on from the clip you already have.

Production-ready content creation across use cases

Built for film, advertising, branding, e-commerce, and gaming — covers on-screen text, brand assets, creative effects, product showcases, UI/UX motion, game visuals, and stylized expression. What those jobs have in common is that the frame contains type, a logo or the product itself, and none of them survive being smeared.

Native stereo sound, made with the picture

H3 generates video with native stereo sound — the audio is not dubbed on afterwards, it comes out of the same pass as the image. Write dialogue, music, ambience and sound effects into the prompt alongside the action and the track follows the shot, which removes the render-then-score-then-sync round trip entirely.

Up to 2K, with the small type still legible

Clips run up to 15 seconds at up to 2K, and 2K is what the model offers by default. It is not a traditional upscaler inventing pixels: the lower-resolution result is sent back through the model together with the original multimodal context and regenerated in context, so captions, logos and interface elements are reconstructed rather than guessed at. That is the same property as “accurate text and brand rendering”, seen from the other side.

Pick the mode before the settings

A2E opens three modes for MiniMax H3. Which one you choose decides what you need to prepare before you write a word of the prompt.

Text to Video

Describe the scene, characters, actions, camera and sound in one prompt to create a complete clip with native stereo audio — no media needed. The fastest way to put an idea on screen when you are starting from nothing.

Image to Video

Animate a single image, or upload first and last frames to control exactly how the sequence starts and ends. Great for posters, product showcases and UI demos. First-and-last frame is not a separate mode here — it lives inside this one, so upload both images when the start and the end both matter.

Reference to Video

Combine image, video and audio references in the same job. Use images for identity, video for motion and camera language, audio for voice and sound. Telling each asset which job it holds is what makes this mode land consistently.

Three steps to your first clip

This is the actual order of operations on A2E. Nothing extra to activate first.

1. Choose mode

Pick text-to-video, image-to-video with first and last frames, or reference-to-video with mixed references. The three are described above; choose wrong and no amount of parameter tuning feels right afterwards.

2. Write the prompt

Describe the scene, motion, and style you want in the video. Sound belongs in the same paragraph — dialogue, music and ambience written next to the action come back in the same generation.

3. Set parameters & generate

Select duration, resolution and aspect ratio, then click generate. The result lands in My Result, where you can carry it on into upscaling, subtitle removal and the rest of the toolbox.

What you can set per generation

These are the controls MiniMax H3 exposes on A2E today — enough to judge whether it clears your delivery bar.

SettingWhat you get
DurationUp to 15 seconds in one clip. The whole arc has to fit inside it, so write the pacing into the prompt instead of hoping the model lands the ending.
Resolution2K, and 2K is the default — there is no separate high-definition tier to remember to switch on.
AudioNative stereo, generated in the same pass as the picture. Nothing to enable, and nothing to sync afterwards.
Input imagesJPG, PNG or WebP, 30 MB or less each, shortest side at least 300 pixels, aspect ratio between 1:2.5 and 2.5:1.
PromptUp to 7,000 characters — room for a shot-by-shot script, not just a sentence.
Number of outputsOne clip per submission.
AccessPaid accounts only.

Five things that save a retry

Give every reference a role

State what each asset controls — character identity, product look, motion, style or voice — to avoid conflicts between inputs. When a mixed-reference job goes wrong, it is usually two assets fighting over the same thing.

Structure longer sequences

Break multi-shot ideas into an ordered progression (e.g. [0-3s]… [3-7s]…) so pacing holds across the whole clip. An unsegmented description tends to be acted out in the first three seconds.

Direct audio with visuals

Describe dialogue, music, ambience and sound effects alongside the action; H3 generates native stereo sound in the same pass. Write only the picture and you have handed the soundtrack to the model.

Define camera language

Specify framing, movement, focus and lens behavior to get intentional cinematography instead of generic motion. “Cinematic” on its own gives the model nothing to hold; a named move does.

Preserve and restrict

List faces, props, text and brand details that must stay consistent, and state what to avoid (hard cuts, morphing, extra text). On work you have to hand over, what must not appear matters as much as what must.

Docs and API

Building MiniMax H3 into your own product? Endpoints, parameters, authentication and billing are in the A2E developer docs, and MCP access is available there too.

Why run MiniMax H3 on A2E

Runs on the account you already have

H3 draws on your existing A2E credits. Nothing to install, no separate signup, no second place to keep track of your renders.

One place for every model

MiniMax, Wan, Seedance, Kling, Sora, Veo, Seedream and more sit behind one account and one balance. Run the same brief through several of them instead of paying for several tools.

Built for delivery

Lip sync, voice cloning, face and head swap, subtitle removal and upscaling live next to generation, so a clip can go from prompt to finished asset without leaving the site.

Related AI video models on A2E

  • Wan 3.0 — Tongyi Wanxiang’s all-in-one video model, with a native 30-second single take. The other half of the pair if your cut needs length rather than mixed references.
  • Seedance 2.5 — ByteDance’s professional multimodal model, also live on A2E. Compare the two on the same brief.
  • Wan 2.6 and Wan 2.6 Flash — an earlier generation, still a fast and economical option.
  • Kling 2.6 and Sora 2 — alternative video models worth running the same prompt through.
  • Lip sync and video to audio — finishing tools for generated footage.