Proofrun protocol / v0.8

Show the work.
Limit the claim.

Proofrun follows one editorial rule: a result should show where it came from, how it was judged, and what the evidence cannot establish.

01Freeze

One prompt, one runtime, one scoring contract.

02Run blind

Model identities stay hidden during judgment.

03Stress

A repeatable probe exercises every artifact.

04Lock

Requirements, scores, and the assessment become immutable.

05Reveal

Only then do model names enter the record.

06Publish

A sanitized, hashed evidence packet crosses the boundary.

01

Frozen brief, declared format

A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.

02

Judge before reveal

Model identities stay hidden until requirement decisions, criterion scores, confidence, and written rationale are locked.

03

Behavior beats the screenshot

A fixed capture records composition. A controlled interaction probe checks whether visible controls respond, state remains coherent, and runtime errors surface.

04

Preserve inconvenient evidence

Wall time, first activity, HTML start, token usage, cost, truncation, browser errors, and probe warnings remain attached to the result.

05

Publish through a boundary

The private Studio keeps credentials, raw responses, and working notes. The public journal receives a verified evidence allowlist. Voxel scenes remain inert data; selected web artifacts open through a separate hash-bound sandbox rather than executing inside the journal.

06

Keep every claim narrow

A Proofrun says what happened here. A result may reveal a useful strength or failure mode, but it does not prove that a model is always better—or always this good.

01

Blind Duel

Two models receive the same frozen test. A winner, tie, confidence label, and scorecard lock before identities are revealed.

02

Solo Run

One model is evaluated against the same explicit requirements and craft criteria without inventing a comparison or opponent.

03

Repeatability study

Several generations repeat one frozen condition. Scores stay in generation order so variability and failed evidence remain visible.

04

Incomplete result

If one or both sides produce no reviewable artifact, Proofrun records the terminal outcome without converting survival into a win. A rerun becomes a new issue.

Playable does not mean trusted by default.

A permanent screenshot remains the canonical visual receipt. When an artifact is also made interactive, its source hash, byte length, selected attempt, publication hash, sandbox profile, and hosted preparation receipt must agree before the journal exposes a launch control.

The journal never executes contestant HTML itself. Web artifacts open on a separate isolated service with a restrictive sandbox. Voxel artifacts are compiled into bounded inert scene data and rendered by version-pinned Proofrun code. If either verification path fails, the preserved screenshot remains available and the interactive launch fails closed.

Interaction is a companion to the record—not a substitute for it.

A report, not a leaderboard.

Models are stochastic, providers change routing, and one prompt samples only one slice of capability. Every result belongs to its date, settings, prompt version, provider response, and preserved artifacts.

Proofrun can support a practical decision. It should not be mistaken for scientific certainty or a permanent ranking. When the evidence is thin, the language stays narrow.

The confidence label describes this verdict—not every possible run.

Historical scores stay historical.

New Evaluation v3 records separate requirement compliance from craft, require every craft score to be deliberately confirmed, map the anchored 1–10 scale from 0 to 100, and preserve a post-lock identity guess or abstention before reveal. Earlier Evaluation v2 reports keep their original score calculation and are not silently rescored.

New runs also preserve a 128-bit SHA-256 commit–reveal proof for contestant order. Earlier runs used a shorter local order seed and did not store a cryptographic pre-review commitment. Their blind reviews remain genuine editorial records, but they should not be read as having the stronger v3 cryptographic guarantee retroactively.

Likewise, the optional AI evidence evaluator’s repeat and swapped-packet figures are a limited smoke check on one synthetic fixture. “Swap divergence” describes disagreement on that fixture; it is not a general measurement of position bias or independent verification.

Read Protocol Notes ↗