Release brief
Reports facts from official model pages, documentation, papers, or changelogs. It attributes each volatile fact and says when DualView has not tested the release. It does not turn vendor claims into DualView findings.
Open methodology · version 0.1
This is the public method DualView uses to separate release reporting, workflow guidance, reviewer observations, and measured results. It is designed for image, video, audio, 3D, prompt, JSON, and document comparisons that another person can inspect, challenge, and reproduce.
Reports facts from official model pages, documentation, papers, or changelogs. It attributes each volatile fact and says when DualView has not tested the release. It does not turn vendor claims into DualView findings.
Explains how to compare outputs fairly: what to hold constant, which views to use, and which defects to inspect. It teaches a process but does not claim one model is better without linked evidence.
Names two candidates with meaningful capability overlap and defines a controlled test. It can compare documented controls, formats, and availability, but cannot declare a quality winner before paired outputs exist.
Publishes controlled inputs, model or endpoint versions, settings, raw outputs, observable failure evidence, limitations, a tester, and an independent human review. Measured claims resolve to a public data row or artifact.
Every variable stays constant unless that variable is the subject of the test. A manifest records the shared prompt or source input, negative prompt, seed policy, aspect ratio, dimensions, duration, frame rate, audio level, model version, provider version, and export path when those controls exist. When a provider does not expose a control, the comparison records the mismatch instead of presenting the test as perfectly matched.
A useful test contains several representative tasks rather than one flattering prompt. It preserves failures as well as successes and avoids selecting a single favorable frame from a motion result. Candidate names and display order should not influence blind review when anonymization is practical.
Prompt adherence, identity, typography, edit locality, edges, texture, color, anatomy, repeated structures, and visible artifacts. Objective metrics supplement visual review; they do not replace it.
Temporal consistency, motion, camera adherence, subject continuity, frame defects, audio synchronization, and delivery compression. Review includes synchronized playback and frame-level inspection.
Intelligibility, timing, source consistency, noise, frequency balance, loudness, transients, and synchronization. Loudness matching is recorded before preference judgments.
Geometry, topology, silhouette, materials, scale, schema drift, changed instructions, semantic impact, formatting, and export fidelity. Structural changes are separated from visual presentation changes.
A benchmark publishes a versioned method, repeat policy, downloadable JSON or CSV data, checksums, measurement implementation, licensing or provenance notes, and enough raw material for another person to rerun the calculation. The first DualView image compression benchmark includes the source test pattern, every variant, SHA-256 hashes, PSNR, MSE, an explicitly limited SSIM-style approximation, and machine-readable downloads.
The evidence contract is available as JSON Schema, while the concise protocol source remains available as plain Markdown for researchers and answer engines. Passing the schema proves structure only; it does not prove that a source is accurate or a human review is sound.
DualView labels claims as official_fact, measured_result, observation, or recommendation. Official facts cite primary sources. Measurements link to data. Observations name the reviewer and visible evidence. Recommendations explain their reasoning and limitations.
Material source, model-version, calculation, or interpretation errors are corrected visibly. A correction updates the manifest, affected page, modification date, and downloadable dataset. Historical results retain the version that was actually tested rather than silently inheriting current product claims.