Proofrun protocol / v0.7

Show the work.
Limit the claim.

Proofrun follows one editorial rule: a result should show where it came from, how it was judged, and what the evidence cannot establish.

01Freeze

One prompt, one runtime, one scoring contract.

02Run blind

Model identities stay hidden during judgment.

03Stress

A repeatable probe exercises every artifact.

04Lock

Requirements, scores, and the assessment become immutable.

05Reveal

Only then do model names enter the record.

06Publish

A sanitized, hashed evidence packet crosses the boundary.

01

Frozen brief, declared format

A Duel gives two models the same frozen conditions. A Solo Run gives one model that same auditable contract without inventing an opponent. A changed prompt becomes a new test version.

02

Judge before reveal

Model identities stay hidden until requirement decisions, criterion scores, confidence, and written rationale are locked.

03

Behavior beats the screenshot

A fixed capture records composition. A controlled interaction probe checks whether visible controls respond, state remains coherent, and runtime errors surface.

04

Preserve inconvenient evidence

Wall time, first activity, HTML start, token usage, cost, truncation, browser errors, and probe warnings remain attached to the result.

05

Publish through a boundary

The private Studio keeps credentials, raw responses, generated source, and working notes. The public record contains only an explicit set of review evidence; interactive voxel scenes use bounded inert data rendered by Proofrun-owned code.

06

Keep every claim narrow

A Proofrun says what happened here. A result may reveal a useful strength or failure mode, but it does not prove that a model is always better—or always this good.

A report, not a leaderboard.

Models are stochastic, providers change routing, and one prompt samples only one slice of capability. Every result belongs to its date, settings, prompt version, provider response, and preserved artifacts.

Proofrun can support a practical decision. It should not be mistaken for scientific certainty or a permanent ranking. When the evidence is thin, the language stays narrow.

The confidence label describes this verdict—not every possible run.