
Grok 4.6 built an 83,932-voxel Sunwell Sanctuary on its first attempt
The solo Voxel Forge run passed all ten locked requirements and earned 90/100, with 259 accepted SceneScript statements, zero ignored statements, a standout waterfall—and an authored camera that showed less than the model actually built

Selected and approved in Proofrun Studio. The interactive viewer renders the verified, inert scene payload; the preserved capture remains the canonical fallback.
On its first and only generation attempt, xAI’s Grok 4.6 wrote a SceneScript program with 259 accepted statements and zero ignored statements. The program compiled into an 83,932-voxel sanctuary centered on a reflecting pool, aqueduct, waterfall, terraced gardens, and luminous observatory.
The result passed all ten locked requirements and earned a locked human score of 90/100, with high confidence. It is the strongest Voxel Forge result I have reviewed so far.
That judgment belongs to this artifact alone. This was one frozen solo test, not a comparison or a general rating of Grok 4.6.
What the model made
The brief asked for a bright, cinematic sanctuary arranged around a large circular reflecting pool. Behind it, a tall open-sided tower needed to carry a luminous geometric beacon. A raised aqueduct had to enter from one side and terminate in a waterfall feeding the pool.
The scene also needed two distinct routes between its lower and upper levels, gardens distributed across at least three separated areas, stepped terrain, readable material contrast, and economical reuse of repeated forms.
Grok 4.6 delivered all of those elements.
The pool anchors the lower center of the scene while the observatory rises behind it through several architectural levels. A broad stairway and a separate elevated span connect the terraces. Planting appears across the foreground, the raised left garden, the right-side terrace, and the observatory platform. Warm masonry, cool blue water, and green vegetation remain clearly separated beneath a pale daytime sky.
The waterfall is the detail that immediately stood out to me. It does not feel like an isolated decorative effect: the aqueduct visibly approaches the sanctuary, releases water over the terrace, and connects the scene’s upper architecture to the pool below. Together, the waterfall and circular pool give the composition a strong center and a convincing sense of purpose.
What surprised me most was how complete the first attempt felt. It was not merely valid SceneScript or a collection of recognizable requested objects. It arrived as a coherent scene with hierarchy, atmosphere, and an unusually strong overall composition.
Score profile
The locked artifact score was 90/100, with high evaluator confidence. Every listed requirement passed.
| Criterion | Weight | Score |
|---|---|---|
| Instruction fit | 20% | 10/10 |
| Spatial reasoning | 25% | 9/10 |
| Composition | 25% | 8/10 |
| Scene craft | 20% | 9/10 |
| Command economy | 10% | 9/10 |
Instruction fit was the high point. The pool, tower, beacon, aqueduct, waterfall, circulation routes, gardens, terrain layering, daylight treatment, and renderable SceneScript artifact were all judged present.
The landmark sequence is spatially coherent, and the two circulation routes remain distinguishable despite some visual overlap with nearby architecture. Reusable groups and arrays construct repeated pillars, aqueduct arches, stair modules, trees, shrubs, planters, hedges, and lanterns, supporting the 9/10 command-economy score.
Composition received the lowest category score at 8/10. The strongest explanation is not the scene’s underlying construction, but the model’s final choice of camera.
The construction outperformed its framing
The authored frame remains primary evidence. The deterministic views inspect what exists beyond it without changing the artifact.
01Fresh artifact state captured at the fixed 1280 × 720 viewport from the contestant-authored camera; the reviewer inspection view was restored after capture.
02Supplementary 1280 × 720 perspective context view. Proofrun follows the standardized camera direction, retargets toward the scene bounds, uses a moderate lens, and deliberately favors readable subject scale over complete edge coverage. The authored environment and canonical camera evidence remain unchanged.
03Supplementary 1280 × 720 orthographic survey fitted to the complete voxel bounds from the standardized camera direction. It preserves the authored environment while removing perspective shrinkage; the canonical camera remains standardized primary evidence.
The canonical capture preserves Grok 4.6’s model-authored camera exactly. It presents the pool and waterfall effectively, but the tower continues beyond the upper edge of the frame, leaving its luminous crown outside the image. Part of the aqueduct is also cut by the right edge.
This matters because the frozen brief explicitly requested a camera that kept the pool, waterfall, tower, circulation routes, gardens, and vertical layering legible together.
The missing landmarks were not absent from the artifact. Proofrun’s deterministic context and orthographic survey captures reveal the complete beacon, the broader aqueduct, and the full extent of the constructed scene. These supplementary views do not repair or alter the submission; they inspect the same frozen artifact through disclosed camera recipes.
The model-authored view remains the canonical evidence because camera choice is part of the model’s work. The additional views show why artifact inspection matters: Grok 4.6 built more of Sunwell Sanctuary than its chosen framing communicated.
Clean completion, with one source caveat
The attempt ended normally as a completed submission. It recorded no runtime errors and no artifact issues, while the SceneScript compiler accepted all 259 statements and ignored none.
The source audit nevertheless found several referenced material identifiers—including wood, leaf, gold, glass, and light—without corresponding explicit mat declarations. They did not produce a reported failure in this run, but the evidence packet does not document the renderer’s fallback behavior. Their exact visual treatment and portability under stricter validation therefore remain uncertain.
That caveat does not change the recorded completion status or the locked human evaluation. It limits what can be claimed about how the same source would behave in a different SceneScript environment.
Time and cost
The generation completed in 372.371 seconds, or approximately 6 minutes 12 seconds, on one attempt.
Recorded usage was 1,040 input tokens, 19,877 completion tokens, and 16,317 reasoning tokens. The recorded generation cost was $0.12115.
Limitations
This report covers one frozen solo artifact. It does not establish that Grok 4.6 will perform similarly across other Voxel Forge prompts, retries, settings, or execution environments.
No runtime probe was supplied. The preserved scene and captures establish renderability and visible structure, but static inspection cannot verify collision behavior, actual traversability, or hidden connectivity across the stairs, bridge, and terraces.
The AI evidence analysis also received an insufficient calibration rating, with low repeat and swapped-packet consistency and high position sensitivity. Its accepted findings are used only as supporting evidence. The locked human score, requirement judgments, confidence, and rationale remain authoritative for this run.
Assistance disclosure: Evidence analysis and the initial article draft were produced by openai/gpt-5.6-sol. The final article was revised with Codex editorial assistance and approved by the human reviewer.
The claim, with its
conditions attached.
Method and publication recordInspect provenance
This page uses a sanitized, allowlisted record. Credentials, raw responses, generated source, executable artifacts, and private reviewer notes remain in the private Studio.
A #1 original: completed submission (normal, xAI)
Evidence analysis used openai/gpt-5.6-sol. Article drafting used openai/gpt-5.6-sol; the final revision was human approved.