# Draft DualView Open Comparison Protocol 0.1

Canonical product guide: https://www.dualview.ai/ai-output-comparison/

This draft defines the evidence DualView intends to require before describing an AI-output comparison as tested. The manifest contract is experimental and is not yet a claim that every existing article or benchmark satisfies it. Its purpose is to keep release reporting, workflow advice, observations, and measured results from being presented as the same kind of claim.

## Publication Types

### Release brief

A release brief reports facts from an official announcement, model card, documentation page, paper, or changelog. It must say that DualView has not tested the model when no controlled outputs exist. It cannot declare a winner, recommend the model, or make comparative performance claims.

### Workflow guide

A workflow guide teaches a reproducible inspection method. It may explain what to hold constant and what defects to inspect, but it cannot claim that one product performs better unless a linked tested comparison supports the claim.

### Comparison brief

A comparison brief names at least two candidates with meaningful capability overlap. It may compare documented controls and publish a future test protocol, but it must disclose when DualView has not generated the paired outputs. It cannot declare a winner or performance advantage until a reviewed evidence manifest links the inputs, outputs, settings, observations, and limitations.

### Tested comparison

A tested comparison includes at least two named candidates, controlled inputs, exact model or endpoint versions, settings, raw outputs, observable failure evidence, limitations, a tester, and a human reviewer. Every measured claim must resolve to a public evidence item.

### Benchmark or dataset

A benchmark adds a versioned method, repeat policy, downloadable data, checksums, measurement implementation, licensing or provenance notes, and enough raw material for another person to reproduce the published result.

## Fairness Controls

Keep every variable constant unless that variable is the subject of the comparison. Record the shared input, prompt, negative prompt, seed policy, aspect ratio, dimensions, duration, frame rate, audio level, model version, provider version, and export path when applicable. If a provider does not expose a control, record that limitation instead of implying the candidates were perfectly matched.

Use several representative tasks rather than a single flattering example. Preserve failed and successful outputs. Do not cherry-pick one frame from a motion result or one crop from a larger image. Separate vendor claims from DualView measurements and named reviewer observations.

## Modality Evidence

- Image: prompt adherence, identity, text rendering, edit locality, edge structure, texture, color, and artifacts.
- Video: temporal consistency, motion, camera adherence, character continuity, frame defects, audio synchronization, and delivery compression.
- Audio: intelligibility, timing, speaker or source consistency, noise, frequency balance, loudness, and synchronization.
- 3D: geometry, topology, silhouette, material, texture, scale, camera consistency, and GLB/GLTF export fidelity.
- Prompt, JSON, and documents: changed instructions, configuration drift, schema differences, semantic impact, and output-format changes.

## Claim Types

- `official_fact`: supported by a named primary source.
- `measured_result`: supported by a public data row, output, metric, or test artifact.
- `observation`: attributed to the named reviewer who inspected the output.
- `recommendation`: a reasoned conclusion whose evidence and limitations are visible.

The draft machine-readable schema is available at https://www.dualview.ai/docs/evidence-manifest.schema.json. Until semantic evidence checks and human approval are fully enforced, it should be treated as a proposed contract rather than a certification badge.

## Corrections

Material source, model-version, calculation, or interpretation errors should be corrected visibly. A correction must update the manifest, the affected page, its modification date, and any downloadable dataset. Historical results should retain the model or endpoint version actually tested instead of silently inheriting current product claims.
