MiniMax H3 Reference-to-Video Guide
A practical guide to using MiniMax H3 reference-to-video with the right asset package, prompt tags, runtime limits, and review loop before you spend more credits.

MiniMax H3 reference-to-video is useful when a single first frame cannot carry the whole instruction. The honest short answer is this: keep the package small, give every reference one job, and test one short clip before you add more motion, more audio, or more style pressure. If you skip that discipline, you are not running a controlled workflow. You are asking the model to negotiate conflicting inputs and then guessing why the result drifted.
Start at the China Video AI homepage if you still need to organize the source material. When you are ready to test, open the MiniMax H3 workspace. The deep link matters because it sends you into the same runtime this guide is describing instead of a generic model picker.
Two supporting posts already cover adjacent jobs. Read How to Make an AI Product Ad From One Image if your problem is still about breaking one still image into short commercial shots. Read the related Seedance 2 Mini reference-to-video checklist if your main goal is a cheaper first pass before you commit to a heavier reference workflow.

The direct answer: one package, one clear responsibility per asset
MiniMax H3 reference-to-video works best when every asset answers a different question. One image can lock the subject. One short video can carry rhythm or camera behavior. One short audio clip can guide timing or sound direction. The prompt then tells the model what must stay faithful and what is free to change.
The common failure is not “MiniMax H3 is random.” The common failure is that the package is overloaded. A user adds several similar images, two contradictory motion clips, a long audio file, and a style-heavy prompt. The output may still look impressive, but the test is no longer diagnostic. You cannot tell which input actually controlled the result.
That is why the first review clip should feel almost conservative. You are not trying to impress a client yet. You are trying to answer three operational questions:
- Did the subject stay coherent?
- Did the motion reference transfer in a useful way?
- Did the package stay small enough that the next iteration is obvious?
If the answer to any of those is “not really,” do not scale the workflow up.
1. Know the real runtime contract before you design the package
The current China Video AI runtime exposes MiniMax H3 as a real supported model, not a placeholder label. It distinguishes plain image-to-video from reference-to-video. Once you move into the reference workflow, the system can accept extra images, short video clips, and short audio clips, but those inputs still need a real reason to exist.
The most practical limits to remember are these:
MiniMax H3 reference-to-video checklist
- output duration: 4 to 15 seconds
- output resolution: 768P or 2K
- reference images: up to 9
- reference videos: up to 3
- reference audios: up to 3
- timed reference length: 2 to 15 seconds each
Those numbers are not a suggestion to max everything out. They are guardrails. The smarter interpretation is that the model allows a fairly rich package, but a rich package should only happen after the smallest useful package is already stable.
If you need a single first frame or a first-and-last-frame pair, you may not need reference-to-video at all. The more inputs you add, the more important the review discipline becomes.
2. Use image-to-video when the shot is already well defined
Reference-to-video is not automatically better than image-to-video. It is just more flexible. That matters because many failed generations are really mode-selection mistakes.
Use plain image-to-video when:
- one first frame already defines the subject clearly;
- the motion is simple enough that the prompt can describe it directly;
- you only need a first frame or a first-and-last-frame transition;
- the real question is whether the scene animates cleanly, not whether several references can be combined.
Use reference-to-video when:
- one image is not enough to preserve the character or product;
- a short clip contains the camera behavior you want to borrow;
- audio timing or rhythm should shape the clip;
- you need one reference for identity and another for motion or scene feel.
This distinction saves credits because it prevents you from carrying unnecessary references into a test that could have been answered with a simpler mode.
3. Build the smallest useful reference package
The strongest MiniMax H3 package is usually smaller than people expect. A good starting package often looks like this:
Reference package v1
- Picture 1: lock the subject or product identity
- Video 1: borrow motion rhythm or camera behavior
- Audio 1: optional, only when timing or pacing matters
That structure is strong because it keeps the jobs separate. If the subject drifts, you look at the identity anchor first. If the camera feels wrong, you inspect the motion clip. If the pacing feels off, you can decide whether the audio reference is helping or confusing the test.
Do not feed several near-duplicate images just because the interface allows it. Extra images only help when each one captures a genuinely different part of the instruction, such as a front view, a side silhouette, or a facial close-up that must survive motion.
The same rule applies to motion references. One short clip with a clear camera move is more useful than three noisy clips that suggest different speeds and different scene geometry.
4. Give each reference one named job in the prompt
MiniMax H3 reference-to-video is easier to control when the prompt names each asset explicitly. The goal is not beautiful prose. The goal is unambiguous instruction.
The public ComfyUI documentation for MiniMax H3 uses position-based reference tags. That is a good practical habit even when you are not working inside ComfyUI, because it forces you to think in responsibilities instead of vibes.
Use <Picture 1> for subject identity and surface details.
Use <Video 1> for camera pace and turn direction.
Use <Audio 1> only for timing accents, not for subject design.
Keep the same subject from <Picture 1>.
Do not copy background objects that conflict with the new scene.
That structure is better than a generic prompt like “make it cinematic and consistent.” The model still needs style language, but style should come after the control language. First define what each asset is doing. Then describe the new scene, framing, motion, and exclusions.
One workable skeleton is:
Subject: keep the same subject and major visual traits from <Picture 1>.
Motion: borrow the camera pace and directional energy from <Video 1>.
Scene: place the subject in a new setting that supports the prompt.
Timing: if <Audio 1> is used, align emphasis with its rhythm only.
Exclude: no duplicated subjects, no random text, no unrelated props, no camera shake.
If one asset does not have a named job, remove it for the first pass.
5. Start with a short review clip, not a final deliverable
A MiniMax H3 reference workflow is easier to debug when the first output is short. Short clips reduce the number of places where identity can drift and make it easier to tell whether the reference logic is working at all.
That is why the first pass should usually answer one narrow question:
- Can the subject stay coherent while motion is introduced?
- Can the camera behavior transfer without breaking the scene?
- Can the audio timing guide the clip without taking over the whole result?
Keep a tiny experiment log while you do this:
Run: h3-r2v-v1
Picture 1: front-reference.png
Video 1: slow-turn.mp4
Audio 1: none
Duration: 6 seconds
Result: subject holds, motion transfers, background becomes too busy in the last second
Next move: simplify background line, keep all references unchanged
This prevents the workflow from collapsing into vague memory. MiniMax H3 can produce striking outputs, but striking is not the same as reproducible.
6. Diagnose the real failure before you add more references
More references are not the default repair. Often they make the failure harder to read.
Use a small diagnosis loop instead:
| Symptom | Likely cause | Smallest next change | | --- | --- | --- | | Subject drifts away from the reference | identity asset is weak or underspecified | replace or strengthen Picture 1 before adding more assets | | Motion feels wrong | the video reference and prompt are describing different moves | keep the same clip and rewrite the motion instruction | | Scene is overloaded | too many references are doing design work | remove the least necessary asset | | Audio makes the result chaotic | audio is guiding too much of the timing | remove Audio 1 for the baseline test | | Output looks technically good but not usable | the package solved style, not task | rewrite the prompt around the real viewer job |
When you review, use the same order every time:
1. subject coherence
2. transfer of motion or timing
3. unwanted objects or duplicate details
4. camera stability
5. whether the clip actually solves the intended task
This order matters because it stops you from praising lighting on a clip that already failed the identity check.
7. MiniMax H3 is strongest when the references are complementary
The best MiniMax H3 packages feel like a team, not a crowd. The image gives the subject. The short motion clip gives behavior. The optional audio gives rhythm. None of them should be fighting for the same job.
That is also why “just add more references” is a poor habit. If two images disagree on pose, or a video implies a different camera speed than the prompt, the model is forced to compromise. Sometimes the compromise still looks interesting. But the result becomes harder to reuse because you no longer know which instruction won.
The practical rule is simple: if you cannot describe the job of each asset in one short sentence, the package is not ready.
8. Decide when MiniMax H3 is the right answer
MiniMax H3 reference-to-video is not a universal default. It is a good answer when you need more control than one frame can provide, especially when you want to combine identity, motion, and timing without switching to a different creative toolchain.
It may be the wrong answer when:
- the shot is already solved by a single first frame;
- the prompt is still vague enough that no reference package can rescue it;
- the task is really a simple product-ad baseline that another lighter workflow can already answer;
- the real need is a different model family or a different creation mode entirely.
Use the Chinese AI video models guide when you need a broader model-family decision. If the project is still closer to a single product-image commercial test, go back to How to Make an AI Product Ad From One Image. If the actual priority is a cheaper early-stage filter, the related Seedance 2 Mini reference-to-video checklist is the better first stop.
9. Review a real motion sample before you expand the package
Watching a real clip changes the review standard. A package that sounds organized on paper can still fail once motion begins. Use the sample below the same way you would use your own first-pass output: inspect coherence, inspect motion transfer, and inspect whether the final frame still looks editable.
When you review, ask:
- does the subject remain readable across the motion arc?
- does the camera move feel intentional rather than rubbery?
- is there a clean frame you could actually keep?
- if the clip failed, do you know which single asset or instruction to change next?
If the answer to the last question is no, the package is too complicated.
10. A production-safe MiniMax H3 checklist
Before you call the workflow ready, run this checklist:
- the mode choice is correct: image-to-video or reference-to-video
- every asset has one named job
- the first pass uses the smallest useful package
- timed references stay short and relevant
- the prompt names each reference explicitly
- the first review clip is short enough to diagnose
- the next change is obvious before another asset is added
This checklist is intentionally narrow. A narrow checklist is what keeps a reference workflow honest. Once the baseline clip passes, then you can justify more duration, more scene ambition, or another controlled reference.
MiniMax H3 is most useful when it helps you preserve what matters and change what should change. The package design is what decides whether that actually happens.
Frequently asked questions
What is MiniMax H3 reference-to-video best for?
Use MiniMax H3 reference-to-video when one prompt is not enough and you need to carry over subject identity, camera behavior, scene style, or sound cues from images, short clips, or audio.
How many references can I use with MiniMax H3 on China Video AI?
The current runtime accepts up to nine reference images, three reference videos, and three reference audios, but the real goal is to keep the package small enough that each asset has one clear job.
Should I use image-to-video or reference-to-video first?
Use image-to-video when one first frame or a first-and-last-frame pair already defines the shot. Use reference-to-video when you need extra images, motion clips, or audio to control the result.
Why does MiniMax H3 ignore one of my references?
The package is often doing too many jobs at once. Reduce the inputs, name each reference explicitly in the prompt, and check whether two assets are giving contradictory instructions.
How long should MiniMax H3 reference clips be?
Keep video and audio references short. The current runtime and public documentation both point to two to fifteen seconds per timed reference, which is enough to carry motion without flooding the generation.
Where can I test MiniMax H3 right now?
Open the MiniMax H3 workspace, upload the smallest useful package, and review one short clip before you expand the test.
The goal is not to build the biggest reference package the interface permits. The goal is to create the smallest package that still answers the reader's real problem. If the package is clear, the next iteration is clear. If the package is muddy, every extra asset makes the workflow more expensive and less teachable.
Go back to the China Video AI homepage if you still need to organize the source material. Go straight to the MiniMax H3 workspace if you already know the subject, the motion sample, and the first review question you need answered. That is the right order: define the task, build the smallest package, then let the model prove it can follow the brief.
