Frozen brief, declared format
A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.
Proofrun follows one editorial rule: a result should show where it came from, how it was judged, and what the evidence cannot establish.
One prompt, one runtime, one scoring contract.
Model identities stay hidden during judgment.
A repeatable probe exercises every artifact.
Requirements, scores, and the assessment become immutable.
Only then do model names enter the record.
A sanitized, hashed evidence packet crosses the boundary.
A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.
Model identities stay hidden until requirement decisions, criterion scores, confidence, and written rationale are locked.
A fixed capture records composition. A controlled interaction probe checks whether visible controls respond, state remains coherent, and runtime errors surface.
Wall time, first activity, HTML start, token usage, cost, truncation, browser errors, and probe warnings remain attached to the result.
The private Studio keeps credentials, raw responses, and working notes. The public journal receives a verified evidence allowlist. Voxel scenes remain inert data; selected web artifacts open through a separate hash-bound sandbox rather than executing inside the journal.
A Proofrun says what happened here. A result may reveal a useful strength or failure mode, but it does not prove that a model is always better—or always this good.
Two models receive the same frozen test. A winner, tie, confidence label, and scorecard lock before identities are revealed.
One model is evaluated against the same explicit requirements and craft criteria without inventing a comparison or opponent.
Several generations repeat one frozen condition. Scores stay in generation order so variability and failed evidence remain visible.
If one or both sides produce no reviewable artifact, Proofrun records the terminal outcome without converting survival into a win. A rerun becomes a new issue.
A permanent screenshot remains the canonical visual receipt. When an artifact is also made interactive, its source hash, byte length, selected attempt, publication hash, sandbox profile, and hosted preparation receipt must agree before the journal exposes a launch control.
The journal never executes contestant HTML itself. Web artifacts open on a separate isolated service with a restrictive sandbox. Voxel artifacts are compiled into bounded inert scene data and rendered by version-pinned Proofrun code. If either verification path fails, the preserved screenshot remains available and the interactive launch fails closed.
Models are stochastic, providers change routing, and one prompt samples only one slice of capability. Every result belongs to its date, settings, prompt version, provider response, and preserved artifacts.
Proofrun can support a practical decision. It should not be mistaken for scientific certainty or a permanent ranking. When the evidence is thin, the language stays narrow.
New Evaluation v3 records separate requirement compliance from craft, require every craft score to be deliberately confirmed, map the anchored 1–10 scale from 0 to 100, and preserve a post-lock identity guess or abstention before reveal. Earlier Evaluation v2 reports keep their original score calculation and are not silently rescored.
New runs also preserve a 128-bit SHA-256 commit–reveal proof for contestant order. Earlier runs used a shorter local order seed and did not store a cryptographic pre-review commitment. Their blind reviews remain genuine editorial records, but they should not be read as having the stronger v3 cryptographic guarantee retroactively.
Likewise, the optional AI evidence evaluator’s repeat and swapped-packet figures are a limited smoke check on one synthetic fixture. “Swap divergence” describes disagreement on that fixture; it is not a general measurement of position bias or independent verification.
Read Protocol Notes ↗