Frozen brief, declared format
A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.
Proofrun follows one editorial rule: a result should show where it came from, how it was judged, and what the evidence cannot establish.
One prompt, one runtime, one scoring contract.
Model identities stay hidden during judgment.
A repeatable probe exercises every artifact.
Requirements, scores, and the assessment become immutable.
Only then do model names enter the record.
A sanitized, hashed evidence packet crosses the boundary.
A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.
Model identities stay hidden until requirement decisions, criterion scores, confidence, and written rationale are locked.
A fixed capture records composition. A controlled interaction probe checks whether visible controls respond, state remains coherent, and runtime errors surface.
Wall time, first activity, HTML start, token usage, cost, truncation, browser errors, and probe warnings remain attached to the result.
The private Studio keeps credentials, raw responses, generated source, and working notes. The public record contains only an explicit set of review evidence; interactive voxel scenes use bounded inert data rendered by Proofrun-owned code.
A Proofrun says what happened here. A result may reveal a useful strength or failure mode, but it does not prove that a model is always better—or always this good.
Models are stochastic, providers change routing, and one prompt samples only one slice of capability. Every result belongs to its date, settings, prompt version, provider response, and preserved artifacts.
Proofrun can support a practical decision. It should not be mistaken for scientific certainty or a permanent ranking. When the evidence is thin, the language stays narrow.