The report
archive.
Every entry is one bounded test with its own prompt, settings, artifacts, failures, and disclosure trail. There is no rolling leaderboard.
claude-fable-5.1 takes on Sunwell Sanctuary
A bright terraced waterscape that tests radial composition, terrain integration, circulation, and daylight readability.
Hy4 preview completed two of three Proofrun tests—and the third stopped at the provider boundary
Across Proofrun’s first Model Release Profile, Tencent’s exact FP8 route scored 65 on Triage Desk and 61 on Verdant Terminal; Weather Story produced no artifact after an HTTP 429, so it remains incident evidence—not a model failure.
Three Attempt Records, One Model Outcome: Qwen3.8 Flash Scores 76 on Triage Desk
After a Proofrun persistence defect and a provider 429 produced no submissions, the only valid artifact passed all ten frozen requirements and earned 76/100—showing why attempt provenance belongs beside model judgment.
qwen3.8-flash takes on Triage Desk
A compact support-operations workspace that tests information architecture, validation, and synchronized multi-step state.
Terra cleared all 10 Triage Desk requirements; its workflow craft scored 79
Proofrun’s first Standard web run produced a complete incident workflow in one attempt. The artifact stayed synchronized through the frozen recipe, while workflow reasoning and information design each scored 7/10.
ox-alpha takes on Orbit Radio
A tactile sci-fi tuner that tests hierarchy, motion, and input handling.
glm-5.3 vs gemini-3.7-flash: Orbit Radio
A tactile sci-fi tuner that tests hierarchy, motion, and input handling.
glm-5.3 and gemini-3.7-flash: an incomplete Orbit Radio run
One model produced a reviewable artifact and the other did not. Proofrun preserves the artifact assessment and the failure record without declaring an automatic head-to-head winner.
gpt-5.6-terra takes on Orbit Radio
A tactile sci-fi tuner that tests hierarchy, motion, and input handling.
grok 4.6 ranged from 8 to 92 across five Prism Garden runs
Five independent generations under one frozen condition produced a median craft score of 86; 4 met every requirement and 1 met none.
Grok 4.6 built an 83,932-voxel Sunwell Sanctuary on its first attempt
The solo Voxel Forge run passed all ten locked requirements and earned 90/100, with 259 accepted SceneScript statements, zero ignored statements, a standout waterfall—and an authored camera that showed less than the model actually built
deepseek-v4-flash-0731 vs hy3: Weather Story
Turns the same small dataset into a useful, visually opinionated forecast.
Fable’s Sky Archive scored 74: a strong beacon undermined by an extra island and unfinished bridges
The solo Voxel Forge run produced valid, economical SceneScript with a clear central landmark, but the locked evaluation penalized its island count, bridge connections, dark palette, and uneven composition.
All Six Requirements Passed, but Prism Garden’s Visual Score Stopped at 6/10
Upstage’s Solar Pro 4 produced a stable, complete generative-art toy that earned 85/100, with functionality outrunning visual distinction in this frozen solo run.
Sol was the visual favorite; Opus 5 led Prism Garden’s scorecard, 77–70
In one blind run, the reviewer preferred Sol’s richer prism work with low confidence, while Opus 5 scored higher on interaction and robustness.
Terra’s steadier physics beat Sonnet 5’s audio edge in Gravity Desk, 87–73
In one blind, frozen comparison, contestant B won with high confidence because stronger physics, control, and presentation outweighed A’s more responsive-feeling sound feedback.