Independent AI model testsIssue 006 / Helsinki

AI models,
judged by the work.

Frozen briefs, blind scoring, working artifacts, costs, failures, and conclusions narrow enough to trust.

How a Proofrun works
Observed, not advertised.

One result.
Three layers of proof.

The page is the argument: working evidence first, human judgment kept explicit, and a publication record readers can inspect for themselves.

01Working artifact

See the thing that was judged.

Proofrun preserves the actual output, its fixed evidence, and—when a safe public renderer exists—the artifact itself.

02Locked judgment

Taste stays visible and bounded.

Requirements and scores lock before identities are revealed. The editor’s view remains human, named, and narrow.

03Public record

The conditions travel with the claim.

Time, cost, attempts, protocol, limitations, and a sanitized record remain attached to every published conclusion.

From frozen brief
to bounded claim.

01Freeze

One prompt, one runtime, one scoring contract.

02Run blind

Model identities stay hidden during judgment.

03Stress

A repeatable probe exercises every artifact.

04Lock

Requirements, scores, and the assessment become immutable.

05Reveal

Only then do model names enter the record.

06Publish

A sanitized, hashed evidence packet crosses the boundary.

Read the methodology
006Sunwell SanctuaryVoxel Forge90/100Published005Weather StoryData visualization92—90Published004Sky ArchiveVoxel Forge74/100Published003Prism GardenCreative coding85/100Published002Prism GardenCreative coding70—77Published001Gravity DeskPhysics toy73—87Published
007Next runNot yet testedOpen·
Browse all reports