Views
No views yet

| Category | Specification |
|---|---|
| Output duration | 4–15 seconds |
| Output aspect ratio | Supports a wide range of aspect ratios, including but not limited to 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 |
| Output resolution | Supports various resolution dimensions. The shorter side is set to 768 pixels by default. 2K | generation can be achieved with H3-Regenerate-2K |
| Output frame rate | 24 FPS |
| Output audio | 32 kHz stereo |
| Supported dialogue languages | Stable support for 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Additional languages are also supported to varying degrees |
| Model Variant | Input Mode | Specifications |
|---|---|---|
| H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images. - No image input: Text-to-video mode - One image input: First-frame-to-video or last-frame-to-video generation - Two image inputs: First-and-last-frame-to-video generation |
| H3-Base-Ref2VA | Omni-reference mode | Supports multi-modal reference inputs: - Images: ≤ 9 images - Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds - Audio: ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds - Mixed inputs: Maximum number of files across all input types is 12 |


<d>, to the tokenizer configuration. When using H3, the tokenizer and associated configuration files provided in the H3 repository are required.1 × 2 × 2 along the (time, height, width) dimensions. As a result, the visual tokens entering the Transformer have an effective spatial downsampling factor of 32×, while the temporal downsampling factor remains 4×.(t, h, w).| Checkpoint | Supported Tasks | Input Conditions | Output | Precision |
|---|---|---|---|---|
| MiniMax-H3 Base FL2VA | Text-to-Audio-Video (t2va), First/Last-Frame-to-Audio-Video (fl2va) | Text; optional first frame, last frame, or both | Video and audio | BF16 |
| MiniMax-H3 Base Ref2VA | Reference-to-Audio-Video (ref2va) | Text with reference images, videos, and/or audio | Video and audio | BF16 |
1<TASK>/
2├── model_index.json
3├── processor/
4├── tokenizer/
5├── text_encoder/
6├── transformer/
7├── visual_vae/
8└── audio_vae/FL2VA/, Ref2VA/) and the diffusers format side by side, so scope the download to what your framework needs:model_index.json is the repository-level public entry. The task-family-specific diffusers indexes remain under FL2VA/model_index.json and Ref2VA/model_index.json.1# Original checkpoint, both task families (SGLang, vLLM):
2hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" "Ref2VA/*" --local-dir MiniMax-H3
3
4# Or a single task family:
5hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" --local-dir MiniMax-H3ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3") fetches exactly the components it needs. See the diffusers documentation for loading recipes.1sglang serve \
2 --model-path MiniMaxAI/MiniMax-H3 \
3 --num-gpus 4 \
4 --ulysses-degree 4 \
5 --performance-mode speed \
6 --host 0.0.0.0 \
7 --port 30010 \
8 --model-variant fl2va1sglang serve \
2 --model-path MiniMaxAI/MiniMax-H3 \
3 --num-gpus 4 \
4 --ulysses-degree 4 \
5 --performance-mode speed \
6 --host 0.0.0.0 \
7 --port 30011 \
8 --model-variant ref2va| Use case | Request | Result |
|---|---|---|
| T2VA | View script | t2va.mp4 |
| FL2VA | View script | fl2va.mp4 |
| Ref2VA | View script | ref2va.mp4 |
1# URL of your SGLang deployment
2SGLANG_DEPLOYMENT_URL="<sglang-deployment-url>"
3
4# MiniMax API endpoint (choose one)
5# CN
6MINIMAX_API_BASE="https://api.minimaxi.com"
7# Global
8# MINIMAX_API_BASE="https://api.minimax.io"
9
10# API token obtained from the MiniMax platform
11TOKEN="<token>"base_video is recommended.| stage | request | result |
|---|---|---|
| H3-Context-IR | View script | json 1{
2 "task": {
3 "id": "<task_id>",
4 "model": "MiniMax-H3",
5 "status": "succeeded",
6 "created_at": "<created_at>",
7 "updated_at": "<updated_at>",
8 "content": {
9 "prompt": "integrated_multimodal_description: [Shot 1] Cinematic, medium wide shot, pushing in slowly. In the cavernous, dimly lit bridge of a starship, sleek metallic consoles with glowing amber displays flank a massive, curved observation window. A female captain, in her late 40s with an athletic build and short silver-streaked black hair, stands in the center midground. She wears a structured, high-collared dark navy military tunic with silver chest insignias. Her back is to the camera, silhouetted against the cool, ambient starlight pouring through the thick glass. She stands perfectly still with her hands clasped tightly behind her back. Outside the window, a massive armada of jagged, dark grey dreadnoughts hovers in tight formation against a deep purple space nebula. The fleet's massive rear thrusters begin to glow with an intense, escalating bright blue light. [Shot 2] At 00:04.500, the camera cuts to a close-up of the captain's face and shakes strongly. The brilliant blue-white light from the fleet's gathering energy reflects vividly in her dark eyes. Suddenly, a blinding white flash floods through the window, completely washing out the background as the fleet jumps to hyperspace. The sheer spatial force violently jolts the bridge, causing the captain from Shot 1 to stagger slightly forward, her shoulders tensing as she visibly braces herself against the physical tremors. As the intense white light fades abruptly, leaving only the dim, empty expanse of the purple nebula reflected on her starkly lit skin, her jaw clenches, and she slowly closes her eyes in the newly emptied space.\noverall_soundscape: A low, resonant hum of the ship's ambient life support systems serves as the baseline, soon drowned out by an audible, escalating, high-pitched electronic whine as the fleet outside charges its hyperdrives. A massive, deafening, bass-heavy boom and sharp crackle erupts during the blinding flash, accompanied by the loud metallic creaking, rattling, and deep thuds of the bridge's bulkheads vibrating under immense physical stress. The intense roaring impact then cuts abruptly back to a hollow, echoing room tone, leaving only the faint, steady hum of the isolated bridge.\nnon_diegetic_music: Cinematic space-opera orchestral score, slow tempo, featuring a solitary, mournful French horn melody over deep, sustained string dissonances that build rapidly in volume and intensity, swelling to a massive orchestral peak before snapping immediately into silence right after the jump."
10 },
11 "duration": 10,
12 "usage": {
13 "total_tokens": 8565,
14 "prompt_tokens": 5650,
15 "completion_tokens": 2915
16 },
17 "ratio": "16:9",
18 "task_type": "h3_context_ir",
19 "modality": "text"
20 }
21} |
| H3-Base | View script | t2va.mp4 |
| H3-Regenerate-2K | View script | t2va_2k.mp4 |
| Reference 2K result by directly calling Open Platform API | View script | h3_direct_2k.mp4 |
| Reference 768P result by directly calling Open Platform API | View script | h3_direct_768p.mp4 |
| stage | request | result |
|---|---|---|
| H3-Context-IR | View script | json 1{
2 "task": {
3 "id": "<task_id>",
4 "model": "MiniMax-H3",
5 "status": "succeeded",
6 "created_at": "<created_at>",
7 "updated_at": "<updated_at>",
8 "content": {
9 "prompt": "For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.\n\nintegrated_multimodal_description: [Shot 1] This is a live-action, cinematic shot with a shallow depth of field. The camera holds a perfectly static shot throughout the entire eight-second duration, capturing a cozy family gathering in a traditional Japanese dining room. The scene opens with a large, intricately patterned blue and white ceramic bowl of ramen in the immediate foreground, rendered in crisp, sharp focus. The bowl sits on a smooth, polished long wooden table. Inside the bowl, a rich, oily golden-brown broth surrounds yellow wavy noodles, topped with two thick, round slices of chashu pork featuring visible fat marbling and a distinct spiral meat pattern. A generous mound of freshly chopped, bright green scallions rests in the center, and a crisp, dark green rectangular sheet of nori seaweed is tucked into the right edge. To the left of the bowl, a pair of light brown wooden chopsticks rests horizontally on a small, dark rectangular chopstick rest, near a small cylindrical ceramic teacup with blue painted patterns. On the right side of the table, a spherical paper lantern with a ribbed bamboo frame sits on a black wooden base. In the background, a large family of seven is gathered around the table, initially appearing as a soft, blurred presence. Behind them, traditional Japanese sliding shoji screens with wooden lattice frames are open, revealing a bright outdoor scene with lush green trees. Early in the clip, the thick, white steam rising from the hot ramen broth immediately intensifies, billowing upwards in thick, swirling clouds that dance continuously above the bowl. As the clip progresses into the middle seconds, the camera maintains its static position while the focus begins a deliberate, smooth shift deeper into the room. The foreground ramen bowl, its vibrant ingredients, and the rising steam gradually soften into a hazy, out-of-focus blur. Simultaneously, the family members in the background come into sharp, detailed clarity. The heavy steam continues to rise from the foreground, creating a dynamic, translucent veil between the camera and the family. With the focus now firmly locked on the background, the vibrant family dinner comes alive. The man in the dark navy blue long-sleeved shirt on the left leans forward, his mouth moving animatedly in a silent exchange. The young girl in the crisp white short-sleeved t-shirt beside him smiles brightly, looking toward the center of the table. The woman on the far left, wearing a soft light blue long-sleeved blouse, turns her head slightly, smiling gently. Across the table, the woman in the light grey button-down shirt smiles broadly, her eyes crinkling, as she rests her hands near her plate. The woman in the dark grey top further back uses her wooden chopsticks to pick up a small piece of food from a central ceramic dish filled with bright red pickled vegetables. The woman in the center back in the light grey sweater smiles gently, her hands clasped softly in front of her, observing the interaction. Throughout the remainder of the clip, the family continues their lively physical interaction, their mouths moving in continuous, silent cadences of conversation, while the thick, white steam from the blurred ramen bowl in the foreground never stops rising, adding a comforting atmosphere to the warm gathering.\n\noverall_soundscape: The soundscape begins with a quiet room tone mixed with the faint, airy rustle of the thick steam billowing from the hot ramen bowl in the foreground, accompanied by the subtle, continuous hissing and bubbling of the rich broth. As the visual focus shifts deeper into the room, the physical sounds of the bustling family dinner become dominant in the foreground. The clear, sharp clinking of ceramic bowls and wooden chopsticks touching plates is clearly heard as the family members reach for food. This is followed by the faint, muffled thud of a cup being set down on the smooth wooden table, and the subtle, rhythmic rustle of cotton and wool clothing as the family members lean forward and gesture, perfectly capturing the lively, physical atmosphere of the shared meal.\n\nnon_diegetic_music: A gentle, heartwarming acoustic guitar melody plays softly in the background, accompanied by the subtle, resonant notes of a traditional Japanese koto. The music maintains a slow, comforting tempo that enhances the cozy, nostalgic, and joyful atmosphere of the family gathering."
10 },
11 "duration": 8,
12 "usage": {
13 "total_tokens": 22822,
14 "prompt_tokens": 12800,
15 "completion_tokens": 10022
16 },
17 "ratio": "16:9",
18 "task_type": "h3_context_ir",
19 "modality": "text"
20 }
21} |
| H3-Base | View script | i2va.mp4 |
| H3-Regenerate-2K | View script | i2va_2k.mp4 |
| Reference 2K result by directly calling Open Platform API | View script | i2va_direct_2k.mp4 |
| Reference 768P result by directly calling Open Platform API | View script | i2va_direct_768p.mp4 |
| stage | request | result |
|---|---|---|
| H3-Context-IR | View script | json 1{
2 "task": {
3 "id": "<task_id>",
4 "model": "MiniMax-H3",
5 "status": "succeeded",
6 "created_at": "<created_at>",
7 "updated_at": "<updated_at>",
8 "content": {
9 "prompt": "subject_definitions:\n<Subject 1> is the young man with short wavy blonde hair, wearing a bright pink suit jacket, matching pink trousers, an unbuttoned white shirt, and silver rings, holding a small black lamb in his arms in <Video 1>.\n<Video 1> is the source video for the editing task.\n<Audio 1> is the synchronized audio track of <Video 1>, providing the background music.\n<Audio 2> is the voice timbre reference for <Subject 1>'s voice, containing a spoken male voiceover.\n\nsummary:\n[video editing + audio reference + audio reuse] The target video is an edited version of <Video 1>. <Subject 1>, wearing a bright pink suit and holding a black lamb, stands in a grassy field with other white lambs in the background. The edit animates <Subject 1>'s face to speak the user-provided dialogue. <Audio 1> is partially reused as the continuous background music, while the target references the calm male voice timbre of <Audio 2> for <Subject 1>'s spoken lines.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - the man retains his identity, wavy blonde hair, pink suit, white shirt, accessories, and the black lamb he holds, with his mouth newly animated to speak.\n<Video 1> (source video editing): fully_preserved - the original camera framing, warm golden hour lighting, grassy hill setting, and background white lambs are maintained while the central character is edited.\n<Audio 1>: partially_copy - the atmospheric background music from <Audio 1> is reused in the target video, mixed beneath the newly added spoken dialogue.\n<Audio 2>: reference - the target audio references the male voice timbre from <Audio 2> to generate <Subject 1>'s spoken dialogue.\n\ndetailed_description:\nThe target video is in realistic photographic style.\n[Shot 1] The shot begins from the source <Video 1>, showing <Subject 1>, a young man with short wavy blonde hair, wearing a bright pink suit jacket, matching pink trousers, and a casually unbuttoned white shirt. He stands confidently in a sunlit green pasture, gently holding a small black lamb securely in his arms. The warm, golden hour lighting casts soft shadows across his face and the bright pink fabric of his suit. Behind him, several white lambs stand and graze on the rolling grassy hill against a clear, pale blue sky. The atmospheric background music from <Audio 1> plays continuously throughout the scene. <Subject 1> physically speaks, his mouth movements naturally syncing to the new dialogue, with his voice timbre referencing the calm male delivery from <Audio 2>. Looking thoughtfully forward, <Subject 1> (S1) speaks softly, <d>[English] Follow the wind, live free.</d> As he delivers the line, he subtly shifts his weight, cradling the resting black lamb while the camera slowly pushes in. <Subject 1> (S1) continues his thought, <d>[English] Leave worries behind, enjoy the moment.</d> Exactly as his voice stops, his lips meet in a relaxed, peaceful smile, and his jaw ceases speaking motion. He then turns his gaze slightly away toward the horizon, gently stroking the black lamb's fleece with his fingers as the camera holds on this tranquil, sunlit state through the end of the video.\n\noverall_soundscape:\nThe soundscape consists of the continuous, atmospheric background music from <Audio 1>, overlaid with the clear, calm male dialogue spoken by the main character, referencing the voice timbre of <Audio 2>.\n\nnon_diegetic_music:\nThe atmospheric, sustained background music from <Audio 1> is reused as the continuous score, playing quietly beneath the spoken dialogue."
10 },
11 "duration": 5,
12 "usage": {
13 "total_tokens": 39299,
14 "prompt_tokens": 33323,
15 "completion_tokens": 5976
16 },
17 "ratio": "16:9",
18 "task_type": "h3_context_ir",
19 "modality": "text"
20 }
21} |
| H3-Base | View script | r2va.mp4 |
| Reference 2K result by directly calling Open Platform API | View script | r2va_2k.mp4 |
| H3 API 2K in Open Platform for reference | View script | r2va_direct_2k.mp4 |
| Reference 768P result by directly calling Open Platform API | View script | r2va_direct_768p.mp4 |