$5 free credits when you sign up Claim now
Whisper Large V3 CT2 now available Test it!
Video Upscaling models now available Test it!
Z-Anime image model Test it!
MiniMax H3 Prompting Guide: How to Write Structured Prompts for Text-to-Video
admin Aug 6, 2026 14 min read

MiniMax H3 Prompting Guide: How to Write Structured Prompts for Text-to-Video

MiniMax H3 generates video and audio from a single prompt. One transformer predicts both – video frames and 32 kHz stereo sound come out of the same denoising process. When a character speaks, their lip movement lands on the exact syllable because the model generated both signals together, not stitched them in post.

That joint generation matters, but so does how you talk to the model. H3 expects a structured document. Write one the wrong way and you will get a fraction of what the model can do.

This guide covers the T2VA (text-to-video-and-audio) mode. If you have a first frame image, the format is almost identical – the only difference is an alignment instruction line at the top.

Why H3 Prompts Look Different

H3 requires a specific structured format. A casual paragraph will generate something, but the model was trained on labelled fields and shot markers, and it performs accordingly.

When you use MiniMax’s hosted product (hailuoai.video or the Open Platform API), your input goes through a preprocessing system called H3-Context-IR. It rewrites your casual prompt into a rigid structured format before the base model sees it. H3-Context-IR is not open-sourced. MiniMax’s own model card says it plainly:

“H3-Context-IR is critical to the quality of the final output, so we strongly recommend incorporating it into your generation pipeline or following the ‘Prompting Guidance’ to build your own context-processing system.”

When you use H3 through an API that runs the base model directly, that preprocessor is gone. You are the Context-IR. You write the structured format yourself, or you accept worse output.

The difference is measurable. We ran the same scenes as casual prompts and structured rewrites on an identical pipeline. The structured versions held tighter compositions, and the generated audio matched what was happening on screen.

The Three Fields

Every H3 prompt has three named sections, written literally in this order:

integrated_multimodal_description: [Shot 1] …

overall_soundscape: …

non_diegetic_music: …

These field names are not labels you can rename or skip. H3 was trained on this format. Leaving them out is like submitting JSON with missing keys.

integrated_multimodal_description

The main body. Everything the viewer sees and most of what they hear goes here: visual style, composition, subject appearance, actions, camera movement, shot changes, spoken dialogue, and any sound that characters can hear (a radio playing, a phone ringing, someone singing).

Write in playback order, from first frame to last. Start [Shot 1] with the visual style and initial composition:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames…

Style options the model responds to well: Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, vintage film.

overall_soundscape

1-4 sentences covering ambient sound, physical action sounds, and non-verbal human sounds across the whole clip. Wind, rain, footsteps, fabric rustling, breathing, impacts.

Do not repeat dialogue or music here. They already have their own homes.

overall_soundscape: Ceramic plates click softly against a wooden counter and cutlery
chimes as it is set down. A kettle simmers and a pan scrapes lightly off the burner,
over quiet household room tone.

non_diegetic_music

1-3 sentences describing score that only the audience hears. Describe instrumentation, tempo, rhythm, and dynamic changes. MiniMax explicitly warns against abstract mood words here – “tense emotional music that builds suspense” will not work. Name the instruments and say what they do.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings
that gradually increase in volume before fading out.

Use N/A when there is no score. Grounded, realistic scenes often sound better without one.

Prompt Length

H3 needs longer prompts than you are used to writing. Short ones produce flat results.

MiniMax’s own Context-IR system outputs roughly 350-450 words in the main description. That is 3-5x longer than a good LTX-2.3 prompt. The model’s text encoder (Qwen3-VL-32B) is built to consume dense, detailed instructions. A 40-word prompt leaves it filling gaps on its own.

Aim for 350-450 words in integrated_multimodal_description for complex scenes, and 150-250 for simpler single-shot clips.

Shot Structure

Cuts and timestamps

[Shot 1] never gets a timestamp. Every later shot gets one in MM:SS.mmm format, strictly increasing:

[Shot 2] At 00:04.500, the camera cuts to…

Decide your total duration first, then place cuts, then write the content. Working backwards from written content produces timestamps that fall outside your clip length.

Practical limits:

  • Budget roughly one cut per 3 seconds. Four shots in 10 seconds is already aggressive.
  • Each shot needs at least 3 seconds to establish anything meaningful.
  • Two-shot scenes work best around 8 seconds. Single beats around 5.

Camera movement

H3 defines camera motion with three dimensions: motion type, amplitude, and speed. Write it as a natural action inside the sentence, not as tags at the end.

✅ The camera pushes in with small amplitude at slow speed toward the folded letter.
❌ …, push in, small amplitude, slow speed.

Available motions: Zoom In/Out, Push In/Pull Out, Pan Left/Right, Truck Left/Right, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, Roll Clockwise/Counterclockwise.

Amplitude (with small amplitude / with large amplitude) and speed (at slow speed / at fast speed) are optional – omit them when you mean medium and normal.

The Dialogue System

H3 generates spoken dialogue with matching lip movement in 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

The grammar is strict.

Speaker IDs. Anyone who speaks gets a stable ID: (S1), (S2). Assigned in order of first vocal appearance. The same person keeps the same ID across every shot. Characters who never speak get no ID.

Dialogue tags. The spoken words go inside <d> tags with a language marker. Everything else – the speaker description, the delivery, the action – stays outside:

The young woman with a quiet, breathy voice (S1) says:
<d>[English] I get off at the next station.</d>

Voiceover. Use the exact phrase says in an off-screen voiceover, then immediately state that the on-screen character’s lips stay closed. Skip the lip-closing line and H3 will animate someone’s mouth to your narration:

The man (S1) says in an off-screen voiceover:
<d>[English] I still remember that road.</d>
while his lips remain completely closed.

Word budget. Natural speech runs about 2.5 words per second. A 10-second clip fits maybe 20-25 spoken words total if you want anything else to happen. Overload the dialogue and H3 either rushes delivery or cuts it off.

One speaker per shot. The single most reliable trick for clean lip sync. If two people need to talk, cut between them.

Everything Must Be Observable

This rule runs through every part of the prompt. Every clause should describe something a viewer can see or hear.

Abstract (will not work)Observable (will work)
She feels abandonedShe lowers her gaze and her shoulders drop
Melancholic atmosphereRain streaks the window, grey light fills the room
Tense, emotional musicA sustained low cello note held under the dialogue
Epic and dramatic sceneThe camera pulls out to reveal the full canyon

The music field works the same way. Compare “creates a sense of wonder” with “a solo French horn at a slow tempo over sustained strings” – only the second gives the audio system something to build.

On-Screen Text

Any text physically visible in frame – signs, banners, labels, phone screens – goes in English double quotation marks, verbatim:

A red neon sign reading “OPEN LATE” glows above the doorway.

Keep on-screen text large and high-contrast. Small text is the first casualty of 8-bit quantization.

Before and After: Why Format Matters

We tested the same scenes written as casual prompts and as structured T2VA rewrites. Every run used 1344×768 at 243 frames (~10 seconds) on identical hardware with a fixed seed. The only variable was prompt format.

Generation time was identical – prompt structure costs encoder tokens, not sampling time.

Product Ad

Casual prompt (64 tokens):

Show coffee beans falling in slow motion onto a rustic wooden table. Transition into a cinematic close-up of hot coffee being poured into a ceramic mug with steam rising naturally. Use warm golden hour lighting streaming through a window. Include smooth tracking shots and shallow depth of field for realism. Mood: comforting, cozy, morning ritual.

Structured rewrite (198 tokens):

integrated_multimodal_description: [Shot 1] Live-action, cinematic commercial, a macro close-up frames roasted coffee beans tumbling onto a rustic wooden table, each bean bouncing and settling into the grain. The camera pulls out with small amplitude at slow speed as the beans come to rest under warm directional light. [Shot 2] At 00:04.500, the camera cuts to a close-up of a white ceramic cup as hot coffee pours in a steady stream, foam swirling on the surface and steam curling upward through the light.

overall_soundscape: Hard coffee beans clatter and skitter across bare wood, settling into a rolling patter. A steady stream of hot liquid pours into a ceramic cup, its pitch rising as the cup fills, over a quiet room tone.

non_diegetic_music: A slow upright-bass line with brushed drums at a relaxed tempo, joined midway by a warm electric piano that fades out at the end.

The casual version gives H3 mood words (“comforting, cozy”) and technique requests (“shallow depth of field”) that it cannot parse into its structured pipeline. The structured version replaces all of that with observable details and precise cut timing. It also populates the two audio fields, which the casual prompt ignores entirely.

AI News Presenter (with dialogue)

Casual prompt (85 tokens):

Generate a professional AI avatar delivering breaking news in a sleek, modern studio. The avatar has realistic facial features, expressive eyes, and subtle hand gestures while speaking. Use a slow camera push-in from a medium shot to a close-up for emphasis. Neutral broadcast lighting illuminates the face evenly, with soft rim lighting to separate the avatar from the background. Include animated ticker graphics at the bottom and a clean, minimalist studio backdrop.

Structured rewrite (207 tokens):

integrated_multimodal_description: [Shot 1] Live-action, broadcast studio look, a medium shot frames a news presenter (S1) behind a dark desk in a sleek modern studio, screens glowing blue behind them. The camera pushes in with small amplitude at slow speed as the presenter looks to camera, and with a clear, measured voice says:
<d>[English] Good evening. We begin tonight with a story developing at this hour.</d> They glance down at the desk, then back up. [Shot 2] At 00:06.500, the camera cuts to a tighter close-up as the presenter continues: <d>[English] Our correspondent is at the scene.</d>

overall_soundscape: A quiet studio room tone with a faint air-handling hum. Paper shifts once on the desk and fabric rustles as the presenter turns.

non_diegetic_music: A short brass-and-synth sting at a moderate tempo under the opening, settling into a low sustained pad.

“Expressive eyes” and “subtle hand gestures” are internal qualities – H3 cannot target them directly. The structured version drops those and instead writes actual dialogue with (S1) speaker IDs, a voice description, and physical actions between lines (“glance down at the desk, then back up”).

Selfie-Style Product Review (UGC with dialogue)

Casual prompt (42 tokens):

Handheld selfie-style video of a person reviewing a product. Natural daylight illuminates face, slight camera shake adds realism. Subtle gestures and facial expressions emphasize opinions. Mood: authentic, relatable, trustworthy.

Structured rewrite (195 tokens):

integrated_multimodal_description: [Shot 1] Live-action, handheld selfie-style, a close-up frames a young woman (S1) holding the camera at arm’s length in a bright room, daylight from a window on her face. The camera shakes slightly at normal speed as she turns a small product toward the lens and, with a warm conversational voice, says: <d>[English] Okay so I’ve been using this for about a week now.</d>She tilts it to show the label. [Shot 2] At 00:06.000, the camera holds as she looks back to lens and adds:
<d>[English] Honestly? I did not expect it to work this fast.</d>

overall_soundscape: A quiet indoor room tone with faint street sound beyond the window. Fabric brushes the microphone as the arm shifts, and the product’s cap clicks once.

non_diegetic_music: N/A

non_diegetic_music: N/A is a deliberate choice. A UGC-style clip with background music would break the authenticity the casual version only asks for by name (“mood: authentic”). Dropping the score, adding handheld shake and room tone, and writing conversational dialogue does more for authenticity than any mood keyword.

Cinematic Cityscape

Casual prompt (36 tokens):

Nighttime futuristic city filled with neon billboards and flying vehicles. Aerial drone-style camera glides between skyscrapers. Wet reflective streets below. Blue-pink lighting palette.

Structured rewrite (190 tokens):

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a wide aerial shot frames a nighttime futuristic city dense with neon billboards in blue and magenta, flying vehicles threading between mirrored skyscrapers. The camera performs a tracking shot forward with large amplitude at slow speed, gliding between towers as reflections streak across wet glass. [Shot 2] At 00:05.000, the camera tilts down with medium amplitude to reveal rain-slick streets far below, neon signage doubled in the standing water as traffic flows through the canyon.

overall_soundscape: A deep city hum carries the rising and falling whine of vehicles passing overhead. Rain hisses against glass and metal, and distant traffic echoes between the towers.

non_diegetic_music: A pulsing synthesizer arpeggio at a steady mid tempo over sustained low bass, with a filtered swell that builds and then drops away.

Four noun phrases (“neon billboards”, “flying vehicles”, “wet streets”, “blue-pink lighting”) describe a still image. The structured version turns that into 10 seconds of camera movement: tracking forward between towers while reflections streak across glass, then tilting down to reveal the rain-slick streets below.

Quick Reference: Common Mistakes

Writing prose instead of fields. A flowing paragraph with no integrated_multimodal_description: label and no [Shot 1] marker will generate, but poorly. H3 was trained on the structured format.

If your prompt is under 100 words, expand it. The Qwen3-VL-32B encoder is built to consume detail, and 40 words leaves it filling gaps on its own.

Mood words in the music field. “Tense and emotional” tells the audio system nothing. Name the instruments, the tempo, and what changes.

Timestamp outside the clip duration. [Shot 2] At 00:09.000 in an 8-second clip. Decide duration first, place cuts second. A related problem: too much dialogue for the duration. Budget 2.5 words per second and leave silence at both ends – 40 words of speech in a 6-second clip will be rushed or truncated.

Sound in the wrong field. A car radio the characters hear is diegetic – it goes in the main description. Rain on a window goes in overall_soundscape. The score goes in non_diegetic_music. Mixing them up degrades output quality.

Voiceover without the lips-closed line. H3 will animate someone’s mouth to your narration unless you explicitly say their lips stay closed.

Tips

  • Decide duration, then place cuts, then write. Working the other way produces out-of-range timestamps.
  • Pin characters by screen position. “On the left”, “in the centre”, “on the right” survives across cuts better than clothing descriptions alone.
  • Reference earlier shots. Write “the captain from Shot 1” or “the young man in the dark-grey hoodie from Shot 1”. H3 needs these anchors to maintain identity.
  • Give the model a repeating event. A rotating lamp, a flickering light, a passing train. It provides a temporal anchor and reduces drift in long takes.
  • Write the soundscape from the picture. Go through your description and ask what each thing sounds like. Feet on gravel, fabric against skin, metal scraping metal. Two extra sentences in overall_soundscape make a noticeable difference.
  • Iterate one field at a time. Fix the seed, change only non_diegetic_music, regenerate. The three fields target different subsystems, so isolating changes helps you find what worked.

Specs at a Glance

ParameterValue
Duration4-15 seconds
FPS24 (fixed)
Resolution768p (short edge); 2K via MiniMax’s hosted module only
Audio32 kHz stereo, generated jointly with video
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Dialogue languages11: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish
Prompt lengthUp to 7,000 characters
Negative promptNot supported
Architecture33B parameter dense omni-modal transformer
Text encoderQwen3-VL-32B

Start Generating

MiniMax H3 is coming to deAPI. When it lands, you will be able to generate video with native audio through the same API you already use for image, voice, and transcription. $5 free credits, no card required.

The prompt format does not change between platforms. A prompt you write today will work the same way when H3 goes live.

→ Get your deAPI key

No subscription No credit card required

Start building with AI in under a minute

Access all models from this article through a single REST API. Start with $5 free credits — no subscription, no credit card.

Migration assistance available talk to an engineer