The report
archive.
Every entry is one bounded test with its own prompt, settings, artifacts, failures, and disclosure trail. There is no rolling leaderboard.
Grok 4.6 built an 83,932-voxel Sunwell Sanctuary on its first attempt
The solo Voxel Forge run passed all ten locked requirements and earned 90/100, with 259 accepted SceneScript statements, zero ignored statements, a standout waterfall—and an authored camera that showed less than the model actually built
deepseek-v4-flash-0731 vs hy3: Weather Story
Turns the same small dataset into a useful, visually opinionated forecast.
Fable’s Sky Archive scored 74: a strong beacon undermined by an extra island and unfinished bridges
The solo Voxel Forge run produced valid, economical SceneScript with a clear central landmark, but the locked evaluation penalized its island count, bridge connections, dark palette, and uneven composition.
All Six Requirements Passed, but Prism Garden’s Visual Score Stopped at 6/10
Upstage’s Solar Pro 4 produced a stable, complete generative-art toy that earned 85/100, with functionality outrunning visual distinction in this frozen solo run.
Sol was the visual favorite; Opus 5 led Prism Garden’s scorecard, 77–70
In one blind run, the reviewer preferred Sol’s richer prism work with low confidence, while Opus 5 scored higher on interaction and robustness.
Terra’s steadier physics beat Sonnet 5’s audio edge in Gravity Desk, 87–73
In one blind, frozen comparison, contestant B won with high confidence because stronger physics, control, and presentation outweighed A’s more responsive-feeling sound feedback.