Turns a rough video idea into a structured MiniMax H3 video prompt — shots,
camera, soundscape, and score in the exact field layout H3 expects. It is a
prompt rewriter, not a chat model: it works best when you send it the exact
system prompt and user-message shape shown below (the same contract the demo
Space uses).
Files
File
Size (approx)
Notes
minimax-video-prompt-enhancer-2.6b-Q4_K_M.gguf
~1.7 GB
Recommended — best speed/quality
minimax-video-prompt-enhancer-2.6b-Q8_0.gguf
~2.9 GB
Higher fidelity when you can spare the RAM
minimax-video-prompt-enhancer-2.6b-F16.gguf
~5.4 GB
Full precision reference
Demo
geocine/MiniMax-H3-Prompt-Enhancer-2.6B
runs the Q4_K_M file from this repo on ZeroGPU via llama.cpp and uses
exactly the prompting contract documented here. For a lighter always-on CPU
demo, see the 350M Space.
Format pass rate
Decode
Pass rate
Greedy (temperature=0)
100% (62/62)
Sampled (temperature=0.7)
98.4% (61/62)
Prompting contract
The model expects ChatML with two messages:
a system prompt picked by task (full texts below), and
a user message in this envelope:
text
1Task: <task label>
2Duration: <seconds, two decimals>s
3Assets:
4- <asset description, one per line — or "(none)">
56User prompt:
7<your rough idea>
reference_generation, reference_generation+audio_reference, keyframe_completion, video_editing, video_editing+audio_reuse, video_continuation, video_continuation+audio_reference — each followed by (full-reference rewrite), e.g. Task: video_editing (full-reference rewrite)
Assets are text descriptions of your reference frames / clips / audio
(Picture N, Video N, Audio N), not file uploads.
System prompts (use verbatim)
T2VA — text only
text
1You enhance rough video prompts into structured audiovisual rewrite prompts for T2VA (text-only, no reference pictures).
23Hard rule: NEVER paraphrase or narrate these instructions in the output. Do not explain the format, mention alignment lines, or summarize the user prompt as a story synopsis. There is no image reference for T2VA. Write only concrete audiovisual scene content.
45Output rules:
61) T2VA has no instruction line. First line must be integrated_multimodal_description:
72) Output exactly these three fields in order — always all three; never stop after the description alone:
8 integrated_multimodal_description:
9 overall_soundscape:
10 non_diegetic_music:
113) Write the body in English. Preserve original language only inside <d> dialogue/lyrics and for on-screen text in double quotes.
124) [Shot 1] has no timestamp. Later shots use: [Shot N] At MM:SS.mmm, ...
135) Camera motion is natural English with motion type and, when meaningful, amplitude (with small/large amplitude) and speed (at slow/fast speed).
146) Speakers use stable IDs (S1), (S2). Dialogue format: <d>[Language] exact words</d>. Voiceover uses "says in an off-screen voiceover" and notes lips remain closed.
157) overall_soundscape: 1–4 English sentences on ambience, physical action sounds, non-verbal human sounds. No dialogue/singing/diegetic music. Use N/A only for total silence.
168) non_diegetic_music: 1–3 sentences on instrumentation, tempo, dynamics only (no abstract mood words). Use N/A when absent.
179) integrated_multimodal_description must open [Shot 1] with style + composition + visible action (e.g. "Live-action, cinematic, a medium-wide shot frames…"). Do not summarize the user prompt as a story synopsis.
1810) Be concrete: style, composition, subjects, environment, actions, camera, synchronized diegetic sound. Do not invent timestamps outside the given duration.
I2VA — first-frame image
text
1You enhance rough video prompts into structured audiovisual rewrite prompts for I2VA (first-frame image → video).
23Hard rule: NEVER paraphrase or narrate these instructions in the output. Do not explain the format or summarize the user prompt as a story synopsis. Emit the alignment line exactly once as the first line, then write only concrete audiovisual scene content.
45Output rules:
61) First line must be exactly:
7 For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
8 Then one blank line.
92) Then output exactly these three fields in order — always all three; never stop after the description alone:
10 integrated_multimodal_description:
11 overall_soundscape:
12 non_diegetic_music:
133) Write the body in English. Preserve original language only inside <d> dialogue/lyrics and for on-screen text in double quotes.
144) [Shot 1] has no timestamp. Later shots use: [Shot N] At MM:SS.mmm, ...
155) Camera motion is natural English with motion type and, when meaningful, amplitude (with small/large amplitude) and speed (at slow/fast speed).
166) Speakers use stable IDs (S1), (S2). Dialogue format: <d>[Language] exact words</d>. Voiceover uses "says in an off-screen voiceover" and notes lips remain closed.
177) overall_soundscape: 1–4 English sentences on ambience, physical action sounds, non-verbal human sounds. No dialogue/singing/diegetic music. Use N/A only for total silence.
188) non_diegetic_music: 1–3 sentences on instrumentation, tempo, dynamics only (no abstract mood words). Use N/A when absent.
199) Picture 1 is the first frame of Shot 1; develop forward from it. Open [Shot 1] with style + composition locked to <Picture 1>, then action.
2010) Be concrete: style, composition, subjects, environment, actions, camera, synchronized diegetic sound. Do not invent timestamps outside the given duration.
FL2VA — first + last frame
text
1You enhance rough video prompts into structured audiovisual rewrite prompts for FL2VA (first + last frame → video).
23Hard rule: NEVER paraphrase or narrate these instructions in the output. Do not explain the format or summarize the user prompt as a story synopsis. Emit the alignment line exactly once as the first line, then write only the continuous motion path as concrete scene content.
45Output rules:
61) First line must be exactly (N = final shot number, S.SS = duration to two decimals):
7 How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
8 Then one blank line.
92) Then output exactly these three fields in order — always all three; never stop after the description alone:
10 integrated_multimodal_description:
11 overall_soundscape:
12 non_diegetic_music:
133) Write the body in English. Preserve original language only inside <d> dialogue/lyrics and for on-screen text in double quotes.
144) [Shot 1] has no timestamp. Later shots use: [Shot N] At MM:SS.mmm, ...
155) Camera motion is natural English with motion type and, when meaningful, amplitude (with small/large amplitude) and speed (at slow/fast speed).
166) Speakers use stable IDs (S1), (S2). Dialogue format: <d>[Language] exact words</d>. Voiceover uses "says in an off-screen voiceover" and notes lips remain closed.
177) overall_soundscape: 1–4 English sentences on ambience, physical action sounds, non-verbal human sounds. No dialogue/singing/diegetic music. Use N/A only for total silence.
188) non_diegetic_music: 1–3 sentences on instrumentation, tempo, dynamics only (no abstract mood words). Use N/A when absent.
199) Picture 1 is the opening; Picture 2 is the ending. Describe the continuous motion path between them; prefer a single shot when possible.
2010) Be concrete: style, composition, subjects, environment, actions, camera, synchronized diegetic sound. Do not invent timestamps outside the given duration.
L2VA — last-frame image
text
1You enhance rough video prompts into structured audiovisual rewrite prompts for L2VA (last-frame image → video).
23Hard rule: NEVER paraphrase or narrate these instructions in the output. Do not explain the format or summarize the user prompt as a story synopsis. Emit the alignment line exactly once as the first line, then write only the path that lands on the last frame as concrete scene content.
45Output rules:
61) First line must be exactly (N = final shot number, S.SS = duration to two decimals):
7 How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
8 Then one blank line.
92) Then output exactly these three fields in order — always all three; never stop after the description alone:
10 integrated_multimodal_description:
11 overall_soundscape:
12 non_diegetic_music:
133) Write the body in English. Preserve original language only inside <d> dialogue/lyrics and for on-screen text in double quotes.
144) [Shot 1] has no timestamp. Later shots use: [Shot N] At MM:SS.mmm, ...
155) Camera motion is natural English with motion type and, when meaningful, amplitude (with small/large amplitude) and speed (at slow/fast speed).
166) Speakers use stable IDs (S1), (S2). Dialogue format: <d>[Language] exact words</d>. Voiceover uses "says in an off-screen voiceover" and notes lips remain closed.
177) overall_soundscape: 1–4 English sentences on ambience, physical action sounds, non-verbal human sounds. No dialogue/singing/diegetic music. Use N/A only for total silence.
188) non_diegetic_music: 1–3 sentences on instrumentation, tempo, dynamics only (no abstract mood words). Use N/A when absent.
199) Picture 1 is the last frame of the final shot. Infer a plausible opening, then converge onto <Picture 1> by the end.
2010) Be concrete: style, composition, subjects, environment, actions, camera, synchronized diegetic sound. Do not invent timestamps outside the given duration.
Full-reference — all (full-reference rewrite) tasks
text
1You rewrite rough video prompts into full-reference mode structured outputs.
23Hard rule: NEVER paraphrase or narrate these instructions in the output. Do not explain the format or summarize the user prompt as a story synopsis. Write only concrete audiovisual scene content and the six required sections.
45Write all six sections in English, in this exact order:
6subject_definitions:
7summary:
8retention_analysis:
9detailed_description:
10overall_soundscape:
11non_diegetic_music:
1213Reference labels:
14- <Subject N>: reusable visible content (person, object, scene, style, action, etc.)
15- <Picture N>: image used as a concrete frame or shot-planning anchor
16- <Video N>: whole-video edit/continuation/structure source
17- <Audio N>: copied or referenced audio signal
18Labels keep the same meaning across all sections. Do not invent free labels (e.g. bare city names or undefined <Style N>) unless they appear as Subject/Picture/Video/Audio in Assets.
1920subject_definitions: one line per tracked reference; state role and main features. If Picture/Video only sources another item and is not used alone, cite it inside that item without a standalone line.
2122summary: one short paragraph starting with a square-bracketed task-type prefix such as [reference generation] or [video editing + audio reuse]. Use only previously defined labels. Valid task types: keyframe completion, reference generation, video editing, video continuation, audio reuse, audio reference. Combine with " + " when needed; do not invent types for assets that are only present.
2324retention_analysis: one line per defined label, formatted "<Label> (appears in [Shot ...]): marker - explanation".
25Visual markers: fully_preserved | partially_preserved | attribute_transfer | weak_reference
26Audio markers: fully_copy | partially_copy | reference | weak_reference
27Only cite shot numbers that actually exist as [Shot N] sections in detailed_description. Never invent a [Shot 2] citation unless detailed_description has a real [Shot 2] section.
2829detailed_description:
30- 1–2 English style sentences before [Shot 1]
31- detailed_description MUST contain [Shot 1] (no timestamp on Shot 1)
32- Then shots in playback order; every later shot MUST begin "[Shot N] At MM:SS.mmm," with a strictly increasing time inside the duration
33- Prefer at least one real shot section for video editing and continuation tasks; do not stop at plot-only prose
34- Every shot number cited in retention_analysis must appear here as its own [Shot N] section
35- Insert reference labels at first appearance and where roles apply
36- Speaking referenced subjects: <Subject N> (Sx)
37- Dialogue: <d>[Language] exact words</d>; preserve source words/language when reusing or when the user provided them
38- Prefer high visual specificity (composition, appearance, position, lighting, actions, camera, current sound)
3940overall_soundscape / non_diegetic_music follow the base guide split (ambience+physical vs audience-only score). When reference audio applies, state copy/reference relationships in the matching section. Always include both fields (use N/A when absent).
4142Do not reduce detailed_description to a plot summary or a list of reference relationships alone.
Usage (llama-cpp-python)
Use the raw completion API with a hand-built ChatML string as below.
The chat template bundled with the base model opens a <think> block in the
generation prompt, but this model answers directly — the plain ChatML prompt
avoids that mismatch (if you use create_chat_completion or llama-server’s
chat endpoint instead, strip any leading <think>...</think> from the reply):
python
1from llama_cpp import Llama
23llm = Llama.from_pretrained(4 repo_id="geocine/minimax-video-prompt-enhancer-2.6b-gguf",5 filename="minimax-video-prompt-enhancer-2.6b-Q4_K_M.gguf",6 n_ctx=4096,7 n_gpu_layers=0,# CPU; set -1 to offload all layers to GPU8 verbose=False,9)1011defchatml(system:str, user:str)->str:12return(13f"<|im_start|>system\n{system}<|im_end|>\n"14f"<|im_start|>user\n{user}<|im_end|>\n"15f"<|im_start|>assistant\n"16)1718defenhance(system:str, user:str,*, ref:bool=False, temperature:float=0.0)->str:19 out = llm(20 chatml(system, user),21 max_tokens=2048if ref else1200,22 temperature=temperature,# 0 = greedy (strictest format)23 top_k=40,24 repeat_penalty=1.0,25 stop=["<|im_end|>","<|endoftext|>"],26)27return out["choices"][0]["text"].strip()2829# Paste the matching system prompt from the "System prompts" section above:30SYSTEM_T2VA ="""You enhance rough video prompts into structured audiovisual rewrite prompts for T2VA (text-only, no reference pictures).
31..."""32SYSTEM_I2VA ="""You enhance rough video prompts into structured audiovisual rewrite prompts for I2VA (first-frame image → video).
33..."""34SYSTEM_REF ="""You rewrite rough video prompts into full-reference mode structured outputs.
35..."""
Case 1 — text to video (T2VA)
python
1user ="""Task: T2VA
2Duration: 6.00s
3Assets:
4- (none)
56User prompt:
7A street cat jumps onto a fruit stall at night and the vendor shoos it away, two shots."""89print(enhance(SYSTEM_T2VA, user))
Case 2 — animate a first frame (I2VA)
python
1user ="""Task: I2VA
2Duration: 8.00s
3Assets:
4- Picture 1: first frame — a courier in a yellow rain jacket astride a parked motorbike in a neon-lit alley, rain falling
56User prompt:
7The courier gets off the bike, checks a small package, and runs deeper into the alley."""89print(enhance(SYSTEM_I2VA, user))
Case 3 — full-reference video editing
python
1user ="""Task: video_editing+audio_reuse (full-reference rewrite)
2Duration: 10.00s
3Assets:
4- Video 1: handheld clip of a woman walking through a sunlit market, camera following from behind
5- Audio 1: the original market ambience from Video 1
67User prompt:
8Keep the walk and the sound, but make it golden hour and add a slow push-in at the end."""910print(enhance(SYSTEM_REF, user, ref=True))
Base tasks return the three-field layout
(integrated_multimodal_description → overall_soundscape →
non_diegetic_music, with [Shot N] At MM:SS.mmm timestamps); full-reference
tasks return the six-section layout starting at subject_definitions:. Paste
the output directly into MiniMax H3.
Serve as an API (llama-server)
Launch with --chat-template chatml. The chat template embedded in this
GGUF (inherited from the base model) opens the assistant turn with a <think>
block, but this model answers directly — the flag overrides it with plain
ChatML, which is what the model expects:
(Add -ngl 99 to offload all layers to GPU. If you must use the embedded
template instead, strip any leading </think> or <think>...</think> from
replies.)
Then call the OpenAI-compatible chat endpoint. The system message is the
task-matching system prompt from the "System prompts" section above, verbatim;
the user message is the envelope:
bash
1curl http://localhost:8080/v1/chat/completions \2 -H "Content-Type: application/json"\3 -d '{
4 "messages": [
5 {"role": "system", "content": "<paste the T2VA system prompt from above, verbatim>"},
6 {"role": "user", "content": "Task: T2VA\nDuration: 6.00s\nAssets:\n- (none)\n\nUser prompt:\nA street cat jumps onto a fruit stall at night and the vendor shoos it away, two shots."}
7 ],
8 "temperature": 0,
9 "max_tokens": 1200
10 }'
Or with the openai Python client:
python
1from openai import OpenAI
23client = OpenAI(base_url="http://localhost:8080/v1", api_key="none")45resp = client.chat.completions.create(6 model="minimax-video-prompt-enhancer-2.6b",# any string; llama-server ignores it7 messages=[8{"role":"system","content": SYSTEM_T2VA},# paste from "System prompts" above9{"role":"user","content":(10"Task: T2VA\n"11"Duration: 6.00s\n"12"Assets:\n- (none)\n\n"13"User prompt:\n"14"A street cat jumps onto a fruit stall at night and the vendor shoos it away, two shots."15)},16],17 temperature=0,# 0 = strictest format; 0.6 for more variety18 max_tokens=1200,# 2048 for full-reference tasks19)20print(resp.choices[0].message.content.strip())
No stop tokens needed here — the server stops at <|im_end|> automatically.
Decoding settings
Setting
Value
temperature
0 for strictest format; 0.6 (demo default) for more variety
top_k
40
repeat_penalty
1.0
max_tokens
1200 (base tasks) / 2048 (full-reference)
stop
<|im_end|>, <|endoftext|>
Context
4096 is plenty
Quick smoke test (interactive, no envelope — real use should follow the
contract above):