Open methodology · version 0.1

Compare outputs without inventing certainty.

This is the public method DualView uses to separate release reporting, workflow guidance, reviewer observations, and measured results. It is designed for image, video, audio, 3D, prompt, JSON, and document comparisons that another person can inspect, challenge, and reproduce.

Evidence rule: a comparison may name a winner only when the tested candidates, controlled inputs, exact versions, settings, raw outputs, measurements, reviewer observations, and limitations are publicly connected. Otherwise it remains a comparison brief or workflow guide.

Four publication levels

Release brief

Reports facts from official model pages, documentation, papers, or changelogs. It attributes each volatile fact and says when DualView has not tested the release. It does not turn vendor claims into DualView findings.

Workflow guide

Explains how to compare outputs fairly: what to hold constant, which views to use, and which defects to inspect. It teaches a process but does not claim one model is better without linked evidence.

Comparison brief

Names two candidates with meaningful capability overlap and defines a controlled test. It can compare documented controls, formats, and availability, but cannot declare a quality winner before paired outputs exist.

Tested comparison

Publishes controlled inputs, model or endpoint versions, settings, raw outputs, observable failure evidence, limitations, a tester, and an independent human review. Measured claims resolve to a public data row or artifact.

Fairness controls

Every variable stays constant unless that variable is the subject of the test. A manifest records the shared prompt or source input, negative prompt, seed policy, aspect ratio, dimensions, duration, frame rate, audio level, model version, provider version, and export path when those controls exist. When a provider does not expose a control, the comparison records the mismatch instead of presenting the test as perfectly matched.

A useful test contains several representative tasks rather than one flattering prompt. It preserves failures as well as successes and avoids selecting a single favorable frame from a motion result. Candidate names and display order should not influence blind review when anonymization is practical.

What gets inspected by modality

Image

Prompt adherence, identity, typography, edit locality, edges, texture, color, anatomy, repeated structures, and visible artifacts. Objective metrics supplement visual review; they do not replace it.

Video

Temporal consistency, motion, camera adherence, subject continuity, frame defects, audio synchronization, and delivery compression. Review includes synchronized playback and frame-level inspection.

Audio

Intelligibility, timing, source consistency, noise, frequency balance, loudness, transients, and synchronization. Loudness matching is recorded before preference judgments.

3D, prompts, and documents

Geometry, topology, silhouette, materials, scale, schema drift, changed instructions, semantic impact, formatting, and export fidelity. Structural changes are separated from visual presentation changes.

Data and reproducibility

A benchmark publishes a versioned method, repeat policy, downloadable JSON or CSV data, checksums, measurement implementation, licensing or provenance notes, and enough raw material for another person to rerun the calculation. The first DualView image compression benchmark includes the source test pattern, every variant, SHA-256 hashes, PSNR, MSE, an explicitly limited SSIM-style approximation, and machine-readable downloads.

The evidence contract is available as JSON Schema, while the concise protocol source remains available as plain Markdown for researchers and answer engines. Passing the schema proves structure only; it does not prove that a source is accurate or a human review is sound.

Claims, review, and corrections

DualView labels claims as official_fact, measured_result, observation, or recommendation. Official facts cite primary sources. Measurements link to data. Observations name the reviewer and visible evidence. Recommendations explain their reasoning and limitations.

Material source, model-version, calculation, or interpretation errors are corrected visibly. A correction updates the manifest, affected page, modification date, and downloadable dataset. Historical results retain the version that was actually tested rather than silently inheriting current product claims.