# AI image and video comparison report template

Use this worksheet to document your own review of two outputs. It does not generate media, calculate a model ranking, or imply that DualView tested the models. Keep private prompts and source media out of public reports unless you intend to share them.

## Question and candidates

- Review date:
- Reviewer:
- Use case and decision to make:
- Modality: image / video
- Model A: provider, exact version, endpoint and availability (public / preview / research)
- Model B: provider, exact version, endpoint and availability
- Official documentation and pricing URLs, with access dates:

## Controlled conditions

- Exact prompt and negative prompt, where supported:
- Input/reference media and permission to publish it:
- Resolution and aspect ratio:
- For video: duration, frame rate and native audio setting:
- Controls supported by each model (reference images, first/last frames, masks, camera or motion controls):
- Seeds, if exposed. The same seed across different models does not imply equivalent initial conditions.
- Number of attempts per candidate; include failures and rejected outputs:
- Any editing, upscaling, compression or transcoding after generation:
- Costs: provider, currency, billing unit, quantity, discounts, retries and total. Do not equate a subscription credit with a dollar without its conversion assumptions.
- Disclose any settings that could not be matched. Separate native-output comparisons from comparisons at normalized delivery settings.

## Inspect the outputs

Open https://www.dualview.ai/compare-two-images/ for images or https://www.dualview.ai/video-comparison/ for video. Load your two files, align them, and inspect the full outputs. Existing examples illustrate the review interface; they are not a controlled benchmark of every model.

| Criterion | Observation for A | Observation for B | Evidence location | Impact on this use case |
| --- | --- | --- | --- | --- |
| Requested subject and composition | | | crop / timestamp | |
| Reference identity and geometry | | | crop / timestamp | |
| Text and small details | | | crop / timestamp | |
| Background, lighting and edges | | | crop / timestamp | |
| Video only: motion and temporal consistency | | | start–end timestamps | |
| Video only: audio or lip sync, when supported | | | start–end timestamps | |
| Editing task only: preserved unedited regions | | | crop / timestamp | |

Record observable defects before assigning a preference. A stylistic preference is not an objective quality measurement. Pixel metrics only make sense for a defined reference and task; a higher score does not establish that one generative model is universally better.

## Evidence and limitations

- Link to the original outputs and settings, if publishable:
- Public DualView comparison link or exported comparison, if publishable:
- Observations directly inspected by the reviewer:
- Claims taken only from provider documentation:
- Independent published evidence and its methodology:
- Missing controls, unavailable models or incomplete tests:
- Selection rule for examples; disclose cherry-picking or convenience sampling:

## Decision

- Preferred candidate for this specific use case, or no defensible preference:
- Evidence supporting that choice:
- Conditions under which the other candidate is preferable:
- What must be tested before production use:

A small sample supports a limited case study, not a general model leaderboard. Label vendor claims, reviewer observations and measured results separately.
