All published model tests

The report
archive.

Every entry is one bounded test with its own prompt, settings, artifacts, failures, and disclosure trail. There is no rolling leaderboard.

001

claude-fable-5.1 takes on Sunwell Sanctuary

A bright terraced waterscape that tests radial composition, terrain integration, circulation, and daylight readability.

Modelanthropic/claude-fable-5.1
Artifact score77/100
Scopestandard
StatusPublished
002

Hy4 preview completed two of three Proofrun tests—and the third stopped at the provider boundary

Across Proofrun’s first Model Release Profile, Tencent’s exact FP8 route scored 65 on Triage Desk and 61 on Verdant Terminal; Weather Story produced no artifact after an HTTP 429, so it remains incident evidence—not a model failure.

FormatModel Release Profile
Slot outcomes65 · Incident · 61
ScopeThree independent Solo runs
StatusPublished
003

Three Attempt Records, One Model Outcome: Qwen3.8 Flash Scores 76 on Triage Desk

After a Proofrun persistence defect and a provider 429 produced no submissions, the only valid artifact passed all ten frozen requirements and earned 76/100—showing why attempt provenance belongs beside model judgment.

Modelqwen/qwen3.8-flash
Artifact score76/100
Scopestandard
StatusPublished
004

qwen3.8-flash takes on Triage Desk

A compact support-operations workspace that tests information architecture, validation, and synchronized multi-step state.

Modelqwen/qwen3.8-flash
Artifact score19/100
Scopestandard
StatusPublished
005

Terra cleared all 10 Triage Desk requirements; its workflow craft scored 79

Proofrun’s first Standard web run produced a complete incident workflow in one attempt. The artifact stayed synchronized through the frozen recipe, while workflow reasoning and information design each scored 7/10.

Modelopenai/gpt-5.6-terra
Artifact score79/100
Scopestandard
StatusPublished
006

ox-alpha takes on Orbit Radio

A tactile sci-fi tuner that tests hierarchy, motion, and input handling.

Modelstealth/ox-alpha
Artifact score56/100
Scopequick
StatusPublished
007

glm-5.3 vs gemini-3.7-flash: Orbit Radio

A tactile sci-fi tuner that tests hierarchy, motion, and input handling.

Winnerz-ai/glm-5.3
Score46—34
Scopequick
StatusPublished
008

glm-5.3 and gemini-3.7-flash: an incomplete Orbit Radio run

One model produced a reviewable artifact and the other did not. Proofrun preserves the artifact assessment and the failure record without declaring an automatic head-to-head winner.

Winner
Scorenull—11
Scopequick
StatusPublished
009

gpt-5.6-terra takes on Orbit Radio

A tactile sci-fi tuner that tests hierarchy, motion, and input handling.

Modelopenai/gpt-5.6-terra
Artifact score78/100
Scopequick
StatusPublished
010

grok 4.6 ranged from 8 to 92 across five Prism Garden runs

Five independent generations under one frozen condition produced a median craft score of 86; 4 met every requirement and 1 met none.

FormatRepeatability study
Generation scores86 · 8 · 92 · 77 · 89
ScopeOne frozen condition
StatusPublished
011

Grok 4.6 built an 83,932-voxel Sunwell Sanctuary on its first attempt

The solo Voxel Forge run passed all ten locked requirements and earned 90/100, with 259 accepted SceneScript statements, zero ignored statements, a standout waterfall—and an authored camera that showed less than the model actually built

Modelx-ai/grok-4.6
Artifact score90/100
Scopestandard
StatusPublished
012

deepseek-v4-flash-0731 vs hy3: Weather Story

Turns the same small dataset into a useful, visually opinionated forecast.

Winnerdeepseek/deepseek-v4-flash-0731
Score92—90
Scopequick
StatusPublished
013

Fable’s Sky Archive scored 74: a strong beacon undermined by an extra island and unfinished bridges

The solo Voxel Forge run produced valid, economical SceneScript with a clear central landmark, but the locked evaluation penalized its island count, bridge connections, dark palette, and uneven composition.

Modelanthropic/claude-fable-5
Artifact score74/100
Scopestandard
StatusPublished
014

All Six Requirements Passed, but Prism Garden’s Visual Score Stopped at 6/10

Upstage’s Solar Pro 4 produced a stable, complete generative-art toy that earned 85/100, with functionality outrunning visual distinction in this frozen solo run.

Modelupstage/solar-pro4
Artifact score85/100
Scopequick
StatusPublished
015

Sol was the visual favorite; Opus 5 led Prism Garden’s scorecard, 77–70

In one blind run, the reviewer preferred Sol’s richer prism work with low confidence, while Opus 5 scored higher on interaction and robustness.

OutcomeSplit decision
Score70—77
Scopequick
StatusPublished
016

Terra’s steadier physics beat Sonnet 5’s audio edge in Gravity Desk, 87–73

In one blind, frozen comparison, contestant B won with high confidence because stronger physics, control, and presentation outweighed A’s more responsive-feeling sound feedback.

Winneropenai/gpt-5.6-terra
Score73—87
Scopequick
StatusPublished