MiniMax H3 Max vs H3: Video API Controls and Dated Costs

By DualView Editorial Team · Published · Sources checked September 25, 2026

Research-based comparison

H3 Max is fal Research's post-trained variant of MiniMax H3. Their hosted text-to-video routes share a useful comparison lane, but differ in output modes and service details. This September 25 review separates documented settings and calculated charges from untested claims about speed or visual quality. DualView has not generated a head-to-head sample set for this article.

Choosing between a base video model and a post-trained variant requires more than comparing names. A different output setting can change both the delivered image and the bill, while prompt rewriting can change the instruction itself. Before deciding whether to migrate, establish the exact route, resolution, duration and expansion behavior that your application will use. Those details determine whether a comparison measures the model difference you intended.

This article examines minimax/h3-max/text-to-video and minimax/h3/text-to-video on fal as of September 25, 2026. fal's H3 Max announcement is dated August 27, so this is a fresh comparison of established endpoints rather than a new-release announcement. The focus is short text-driven video requests. Image-conditioned and multimodal-reference routes need their own comparison; their capabilities are not silently attributed to these two text routes.

Sources: H3 Max text-to-video endpoint and dated rates · H3 text-to-video endpoint and rates · H3 Max text-to-video input and output schema · H3 text-to-video input and output schema · Introducing H3 Max by fal, August 27, 2026 · MiniMax H3 family overview on fal

Documented model comparison

Provider descriptions and evidence gaps as of September 25, 2026
Integration questionH3 Max text-to-videoH3 text-to-video
Variant identityPost-trained by fal Research from H3Base MiniMax H3 route
Resolution modes480P, 768P, 1080P; 1080P uses latent refinement480P, 768P, 2K, 4K; higher modes upscale 768P
Default output and duration768P; five seconds2K; five seconds
Prompt controlExpansion mode required; seed exposedExpansion mode optional; seed exposed
Audio inputOptional target_audio_urlOptional target_audio_url
DeliveryHosted API, MP4 file outputHosted API, MP4 file output
Weights evidenceReviewed sources do not establish downloadable Max weightsFamily overview describes open weights; inspect license separately

Sources: H3 Max text-to-video endpoint and dated rates · H3 text-to-video endpoint and rates · H3 Max text-to-video input and output schema · H3 text-to-video input and output schema · Introducing H3 Max by fal, August 27, 2026 · MiniMax H3 family overview on fal

Match native output before comparing delivered pixels

The schemas provide a concrete starting point: both offer 480P and 768P, while their higher modes follow different paths. For an initial comparison, explicitly request a common native mode instead of accepting each route's default. Then make a separate comparison of delivery modes if a client needs a larger final frame. A larger file is a deliverable property, not evidence that motion or fine detail was generated at that scale.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema

For review, keep the original files and record actual frame dimensions and duration. Inspect small lettering, thin edges and texture changes during movement, rather than relying on a single appealing frame. If you resize clips to the same viewing area, retain a way to inspect each original. Otherwise, your viewing transform can conceal the very reconstruction differences you are trying to assess. These are proposed checks, not observations from H3 output generated by DualView.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema

Treat prompt expansion as part of the experiment

Both schemas expose a seed and prompt expansion, but the fields are not identical: Max marks expansion as required, whereas the base route does not. The base documentation additionally describes a fast expansion option. For a controlled first pass, set expansion deliberately on both requests and retain any returned expanded text. Do not assume that choosing the same seed creates corresponding frames across different model variants.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema

There are two legitimate questions here. One is how each video model responds to the same authored instruction; the other is how each complete service turns a brief into usable footage. Keep those evaluations separate. An expansion system might add a camera move, change shot structure or emphasize a different subject detail. When that happens, inspect the returned instruction before attributing a visual difference solely to post-training. If the effective instruction is unavailable, document that evidence gap in the result.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema

Keep generated sound separate from supplied sound

The H3 family overview describes native stereo sound alongside video. The reviewed text-to-video schemas also accept target_audio_url, but describe a specific operation: the original recording replaces the output soundtrack, with trimming or silence padding as needed. That control should not be presented as proof of voice transfer or generated dialogue quality. Audio is discussed here only as a feature of the video models.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema · MiniMax H3 family overview on fal

For a dialogue-oriented test, record whether a soundtrack was generated or supplied. Listen for event timing while watching the same frames, and keep any externally edited audio in a separately named file. A replaced soundtrack can make a clip feel more finished without establishing that the model synchronized a newly generated voice. For silent visual review, mute both clips consistently; for audiovisual review, preserve their original audio treatment. This makes the eventual conclusion specific enough to guide a real production choice.

Sources: H3 Max text-to-video input and output schema · H3 text-to-video input and output schema · MiniMax H3 family overview on fal

Budget at the same resolution and watch the discount date

On September 25, fal lists H3 Max at $0.025 per second for 480P, $0.04 for 768P and $0.08 for 1080P. Its endpoint calls these promotional rates, with the discount ending September 30 and subsequent rates of $0.05, $0.08 and $0.16 respectively. The older announcement referred to a first-week offer; the dated endpoint notice is the basis for this estimate, not an assumption that the original offer stayed unchanged.

Sources: H3 Max text-to-video endpoint and dated rates · H3 text-to-video endpoint and rates · Introducing H3 Max by fal, August 27, 2026

The base H3 page lists $0.05 per second for 480P, $0.06 for 768P, $0.13 for 2K and $0.16 for 4K. A five-second 768P request therefore totals $0.20 on Max at the listed promotion versus $0.30 on base H3. At Max's stated subsequent 768P rate, that same calculation becomes $0.40. These are duration-times-rate estimates for one output, excluding retries, taxes and other workflow costs; they establish no quality equivalence.

Sources: H3 Max text-to-video endpoint and dated rates · H3 text-to-video endpoint and rates · Introducing H3 Max by fal, August 27, 2026

Separate service timing from a speed headline

fal's announcement attributes improvements to both post-training and its inference system. Those are vendor claims, not measurements performed for this article. Max's output schema defines the inference timing field as GPU denoising time and notes that it may be absent on some routes. That number cannot, by itself, describe a user's wait from submitting a request to having a playable file.

Sources: H3 Max text-to-video input and output schema · Introducing H3 Max by fal, August 27, 2026

An application deciding between the routes should retain submission, queue, completion and file-availability timestamps in addition to any backend timing. Use the same request class and repeat it across more than one time window. Report failures and queue delays alongside completed requests. A responsive creative preview and a large overnight batch have different priorities, so a single vendor timing example cannot settle both decisions. No request-time measurements, preference rankings or independent leaderboard conclusions are asserted here.

Sources: H3 Max text-to-video input and output schema · Introducing H3 Max by fal, August 27, 2026

Choose a route by the constraint you actually have

For a team shipping through fal, the immediate decision is which endpoint fits the required output mode, controls and expected request budget. A team that needs self-hosting has a different investigation: fal describes base H3 as open-weight, while the reviewed Max sources do not establish downloadable variant weights. That distinction does not supply license terms, hardware costs or proof that a local deployment behaves like the hosted service.

Sources: MiniMax H3 family overview on fal · Introducing H3 Max by fal, August 27, 2026

Before a production switch, assemble a small retained test set around your own scenes: subject continuity, action completion, legible detail and soundtrack timing. Define rejection criteria first and preserve failed attempts. DualView can help inspect matching moments side by side, but the comparison view cannot turn undocumented settings into evidence. Choose the route whose documented constraints fit the job, then base any quality conclusion on retained outputs and a stated review method. Keep promotional-budget decisions dated so they can be revisited.

Sources: MiniMax H3 family overview on fal · Introducing H3 Max by fal, August 27, 2026

Open DualView video comparison to review your own video outputs.

Frequently asked questions

Does a 4K output mean the base model generated every detail at 4K?

No. The reviewed H3 schema explicitly describes its 2K and 4K modes as upscaling a 768P base result. Evaluate the delivered file separately from the native generation mode.

Does the shared five-second default establish every allowed duration?

No. A default value is not a complete range guarantee. The base endpoint description mentions five to fifteen seconds; verify constraints for the exact Max route before extending a request.

Has DualView measured which H3 variant produces better video?

No. This article reports documented interface differences and arithmetic from dated provider rates. It includes no generated comparison samples, measured request times or empirical visual-quality verdict.

Official sources

No paid model generations were performed for this article. The proposed review procedure is editorial analysis, not an executed experiment.

About the author

DualView Editorial Team documents image and video model capabilities and reproducible comparison methods. Read about the editorial team.