MiniMax H3 Max vs H3: Video API Controls and Dated Costs
H3 Max is fal Research's post-trained variant of MiniMax H3. Their hosted text-to-video routes share a useful comparison lane, but differ in output modes and service details. This September 25 review separates documented settings and calculated charges from untested claims about speed or visual quality. DualView has not generated a head-to-head sample set for this article.
Choosing between a base video model and a post-trained variant requires more than comparing names. A different output setting can change both the delivered image and the bill, while prompt rewriting can change the instruction itself. Before deciding whether to migrate, establish the exact route, resolution, duration and expansion behavior that your application will use. Those details determine whether a comparison measures the model difference you intended.
This article examines minimax/h3-max/text-to-video and minimax/h3/text-to-video on fal as of September 25, 2026. fal's H3 Max announcement is dated August 27, so this is a fresh comparison of established endpoints rather than a new-release announcement. The focus is short text-driven video requests. Image-conditioned and multimodal-reference routes need their own comparison; their capabilities are not silently attributed to these two text routes.
Documented model comparison
| Integration question | H3 Max text-to-video | H3 text-to-video |
|---|---|---|
| Variant identity | Post-trained by fal Research from H3 | Base MiniMax H3 route |
| Resolution modes | 480P, 768P, 1080P; 1080P uses latent refinement | 480P, 768P, 2K, 4K; higher modes upscale 768P |
| Default output and duration | 768P; five seconds | 2K; five seconds |
| Prompt control | Expansion mode required; seed exposed | Expansion mode optional; seed exposed |
| Audio input | Optional target_audio_url | Optional target_audio_url |
| Delivery | Hosted API, MP4 file output | Hosted API, MP4 file output |
| Weights evidence | Reviewed sources do not establish downloadable Max weights | Family overview describes open weights; inspect license separately |
Match native output before comparing delivered pixels
The schemas provide a concrete starting point: both offer 480P and 768P, while their higher modes follow different paths. For an initial comparison, explicitly request a common native mode instead of accepting each route's default. Then make a separate comparison of delivery modes if a client needs a larger final frame. A larger file is a deliverable property, not evidence that motion or fine detail was generated at that scale.
For review, keep the original files and record actual frame dimensions and duration. Inspect small lettering, thin edges and texture changes during movement, rather than relying on a single appealing frame. If you resize clips to the same viewing area, retain a way to inspect each original. Otherwise, your viewing transform can conceal the very reconstruction differences you are trying to assess. These are proposed checks, not observations from H3 output generated by DualView.
Treat prompt expansion as part of the experiment
Both schemas expose a seed and prompt expansion, but the fields are not identical: Max marks expansion as required, whereas the base route does not. The base documentation additionally describes a fast expansion option. For a controlled first pass, set expansion deliberately on both requests and retain any returned expanded text. Do not assume that choosing the same seed creates corresponding frames across different model variants.
There are two legitimate questions here. One is how each video model responds to the same authored instruction; the other is how each complete service turns a brief into usable footage. Keep those evaluations separate. An expansion system might add a camera move, change shot structure or emphasize a different subject detail. When that happens, inspect the returned instruction before attributing a visual difference solely to post-training. If the effective instruction is unavailable, document that evidence gap in the result.
Keep generated sound separate from supplied sound
The H3 family overview describes native stereo sound alongside video. The reviewed text-to-video schemas also accept target_audio_url, but describe a specific operation: the original recording replaces the output soundtrack, with trimming or silence padding as needed. That control should not be presented as proof of voice transfer or generated dialogue quality. Audio is discussed here only as a feature of the video models.
For a dialogue-oriented test, record whether a soundtrack was generated or supplied. Listen for event timing while watching the same frames, and keep any externally edited audio in a separately named file. A replaced soundtrack can make a clip feel more finished without establishing that the model synchronized a newly generated voice. For silent visual review, mute both clips consistently; for audiovisual review, preserve their original audio treatment. This makes the eventual conclusion specific enough to guide a real production choice.
Budget at the same resolution and watch the discount date
On September 25, fal lists H3 Max at $0.025 per second for 480P, $0.04 for 768P and $0.08 for 1080P. Its endpoint calls these promotional rates, with the discount ending September 30 and subsequent rates of $0.05, $0.08 and $0.16 respectively. The older announcement referred to a first-week offer; the dated endpoint notice is the basis for this estimate, not an assumption that the original offer stayed unchanged.
The base H3 page lists $0.05 per second for 480P, $0.06 for 768P, $0.13 for 2K and $0.16 for 4K. A five-second 768P request therefore totals $0.20 on Max at the listed promotion versus $0.30 on base H3. At Max's stated subsequent 768P rate, that same calculation becomes $0.40. These are duration-times-rate estimates for one output, excluding retries, taxes and other workflow costs; they establish no quality equivalence.
Separate service timing from a speed headline
fal's announcement attributes improvements to both post-training and its inference system. Those are vendor claims, not measurements performed for this article. Max's output schema defines the inference timing field as GPU denoising time and notes that it may be absent on some routes. That number cannot, by itself, describe a user's wait from submitting a request to having a playable file.
An application deciding between the routes should retain submission, queue, completion and file-availability timestamps in addition to any backend timing. Use the same request class and repeat it across more than one time window. Report failures and queue delays alongside completed requests. A responsive creative preview and a large overnight batch have different priorities, so a single vendor timing example cannot settle both decisions. No request-time measurements, preference rankings or independent leaderboard conclusions are asserted here.
Choose a route by the constraint you actually have
For a team shipping through fal, the immediate decision is which endpoint fits the required output mode, controls and expected request budget. A team that needs self-hosting has a different investigation: fal describes base H3 as open-weight, while the reviewed Max sources do not establish downloadable variant weights. That distinction does not supply license terms, hardware costs or proof that a local deployment behaves like the hosted service.
Before a production switch, assemble a small retained test set around your own scenes: subject continuity, action completion, legible detail and soundtrack timing. Define rejection criteria first and preserve failed attempts. DualView can help inspect matching moments side by side, but the comparison view cannot turn undocumented settings into evidence. Choose the route whose documented constraints fit the job, then base any quality conclusion on retained outputs and a stated review method. Keep promotional-budget decisions dated so they can be revisited.
Open DualView video comparison to review your own video outputs.
Frequently asked questions
Does a 4K output mean the base model generated every detail at 4K?
No. The reviewed H3 schema explicitly describes its 2K and 4K modes as upscaling a 768P base result. Evaluate the delivered file separately from the native generation mode.
Does the shared five-second default establish every allowed duration?
No. A default value is not a complete range guarantee. The base endpoint description mentions five to fifteen seconds; verify constraints for the exact Max route before extending a request.
Has DualView measured which H3 variant produces better video?
No. This article reports documented interface differences and arithmetic from dated provider rates. It includes no generated comparison samples, measured request times or empirical visual-quality verdict.
Official sources
- H3 Max text-to-video endpoint and dated rates
- H3 text-to-video endpoint and rates
- H3 Max text-to-video input and output schema
- H3 text-to-video input and output schema
- Introducing H3 Max by fal, August 27, 2026
- MiniMax H3 family overview on fal
No paid model generations were performed for this article. The proposed review procedure is editorial analysis, not an executed experiment.