Official-source specification comparison · no winner declared
FLUX 3 vs Grok Imagine 1.5
FLUX 3 documents a broad Early Access video system with native audio and up to 20-second generations. Grok Imagine 1.5 documents a focused image-to-video API. Compare them directly only on the shared still-image animation task; treat every other capability as workflow context.
Evidence boundary: DualView has not yet run the controlled benchmark described below. Specifications come from official pages. Vendor samples are not matched for prompt, input, settings, selection, or post-processing.
Official showcase material
These are real videos published by the model providers. Watch them as examples of each vendor's presentation, not as evidence that one model beats the other.
FLUX 3 official showcase
Official BFL launch-page media streamed from the vendor CDN with preloading disabled.
Open the official sourceGrok Imagine 1.5 official showcase
xAI-published launch-page output preserved without re-encoding.
Open the official sourceWhat the official documentation actually supports
Black Forest Labs and xAI frame these releases differently. FLUX 3 spans text, image, video, and audio workflows, including T2V, I2V, V2V, continuation, keyframes, dialogue, and chaining. Grok Imagine 1.5 begins with one source image and aims to animate it faithfully according to camera, pacing, atmosphere, physics, and sound instructions.
BFL reports that FLUX 3 was preferred over Grok Imagine Video in up to 69 percent of its preliminary comparisons. This is a vendor-reported Early Access result; the announcement does not identify it as a DualView test, and the label does not establish that every sample used Grok Imagine 1.5. It is context for designing an independent replication, not a verdict.
A useful replication should publish the starting frames, exact prompts, model identifiers, settings, failed attempts, and every selected output. The first lane should hold duration and resolution equal. A second lane can explore FLUX 3's 20-second range and Grok's current resolution options without mixing those operational differences into the core quality score.
| Question | FLUX 3 | Grok Imagine 1.5 |
|---|---|---|
| Fair overlapping lane | Image-to-video from a starting frame or visual reference | Single-image-to-video |
| Broader generation | T2V, V2V, continuation, keyframes, and chaining documented | Not documented for model 1.5 |
| Duration evidence | Up to 20 seconds | Official launch example uses 10 seconds; model page states no maximum |
| Audio | Native audio on all documented video outputs | Prompt-directed sound design |
| Status | Early Access; evaluations explicitly preliminary | API model with preview and dated aliases in current docs |
| Independent matched result | Not measured by DualView | Not measured by DualView |
Model profiles in one paragraph
FLUX 3
Black Forest Labs documents Early Access video with native audio up to 20 seconds, text-to-video, image-to-video, video-to-video, video-audio continuation, keyframe transitions, multilingual dialogue, and clip chaining.
Grok Imagine 1.5
xAI introduces Grok Imagine 1.5 as an image-to-video model: one still image plus a motion prompt, including camera movement, atmosphere, physics, sound direction, and source-image fidelity.
A fair comparison protocol
The protocol below isolates shared capabilities before exploring model-specific advantages. Save original files and settings so another reviewer can reproduce the result.
Use identical source images and a balanced set of camera, action, atmosphere, and sound prompts.
Fix the first pass at 10 seconds and 720p to align with both vendors' published examples.
Retain all outputs and retries; blind the model labels during human preference scoring.
Report source fidelity, instruction adherence, motion coherence, audio synchronization, latency, and cost separately.
Which one belongs on the shortlist?
Start with Grok Imagine 1.5 when the workflow is intentionally one-image-in and video-out. Test FLUX 3 when the same project also needs text-only generation, V2V, continuation, keyframes, or native-audio breadth. Do not treat BFL's preliminary vendor result as a substitute for your own matched inputs.
What this page does not claim
This FLUX 3 and Grok Imagine 1.5 review does not declare an overall quality winner or turn provider preference figures into independent evidence. It leaves undocumented limits blank. Recheck the linked first-party pages because access, aliases, pricing, and output options can change.
Official sources
- Black Forest Labs FLUX 3 announcement
- xAI Grok Imagine 1.5 announcement
- Grok Imagine 1.5 current model documentation