Repeatability study 007 / Creative coding / Published August 13, 2026

grok 4.6 ranged from 8 to 92 across five Prism Garden runs

Five independent generations under one frozen condition produced a median craft score of 86; 4 met every requirement and 1 met none.

Bounded resultOne frozen condition produced four fully compliant artifacts and one complete submission that met none of the requirements.Five generations expose material variability here. They do not establish a model-wide distribution or rating.
Modelx-ai/grok-4.6TestPrism Garden v2Protocolv0.7StudyRP-MSROO2DW-E03530
Generation scores86 · 8 · 92 · 77 · 89
Median86
Range892
Reported cost$0.4375
01 / Five outcomes

The same brief did not
produce the same result.

Every card is shown in generation order. Four targets completed the frozen seven-bloom recipe and preserved a 1280 × 720 demonstration capture. The 8/100 target retained its failed assertion receipt without a replacement image.

Generation 0186/100
grok 4.6 generation 1: Standardized bloom demonstration.
Seven blooms planted at the frozen Prism Garden v1 coordinates after reset, then captured after the standardized stabilization window.
Compliance
100%
Wall time
2:33
Cost
$0.0624
completed submissionrecipe captured
Generation 028/100
No manufactured image

Demonstration differs from the reset state: The visual fingerprint did not change after the frozen actions.

Compliance
0%
Wall time
3:37
Cost
$0.0944
completed submissionrecipe failed
Generation 0392/100
grok 4.6 generation 3: Standardized bloom demonstration.
Seven blooms planted at the frozen Prism Garden v1 coordinates after reset, then captured after the standardized stabilization window.
Compliance
100%
Wall time
3:34
Cost
$0.0950
completed submissionrecipe captured
Generation 0477/100
grok 4.6 generation 4: Standardized bloom demonstration.
Seven blooms planted at the frozen Prism Garden v1 coordinates after reset, then captured after the standardized stabilization window.
Compliance
100%
Wall time
2:40
Cost
$0.0746
completed submissionrecipe captured
Generation 0589/100
grok 4.6 generation 5: Standardized bloom demonstration.
Seven blooms planted at the frozen Prism Garden v1 coordinates after reset, then captured after the standardized stabilization window.
Compliance
100%
Wall time
3:55
Cost
$0.1111
completed submissionrecipe captured
02 / Score spread

Four passed everything.
One passed nothing.

The five locked craft scores were 86, 8, 92, 77, 89. Their median was 86, while the observed range stretched from 8 to 92.

The low score was not a provider failure or missing artifact. Generation 2 returned a complete HTML submission, but the blind review marked all six frozen requirements failed. Completion and compliance therefore remain separate facts.

Full compliance4/5 generations
Compliance collapse1/5 generations
No-submission failures0/5 generations
GenerationCraftComplianceTimeReported cost
186100%2:33$0.0624
280%3:37$0.0944
392100%3:34$0.0950
477100%2:40$0.0746
589100%3:55$0.1111
03 / Review checks

Two zeros
and a plus six.

The review design included a separate cold rescore before launch and two exact duplicate artifacts hidden inside the nine-item blind queue. These checks describe this reviewer in this session; they are not a universal correction factor.

Cold rescore+6

The rehearsal score was six points higher than the artifact’s original review.

Hidden duplicate 10

BLIND ARTIFACT 01 and BLIND ARTIFACT 05 received the same score.

Hidden duplicate 20

BLIND ARTIFACT 02 and BLIND ARTIFACT 09 received the same score.

What this does not showTwo zero duplicate deltas do not prove perfect reviewer consistency, and the +6 cold rescore should not be subtracted from any generation. The study reports both observations without “correcting” the five locked scores.
04 / Frozen method

One condition,
kept attached.

Model and samplinggrok 4.6High reasoning · T=0.7
Generation limit16,000output tokens · 20m wall limit
Public evidence4 captures1 explicit failed receipt
Study and publication recordInspect provenance

The journal renders only the allowlisted record and hash-bound recipe images. It does not execute any generated contestant HTML. The researcher notebook, queue secret, raw responses, generated source, credentials, router metadata, and private reviewer notes remain in Studio.

Condition hashsha256:da80b854ec27d19f4f6510213108896601e9e7c64603ecd10fe43314d1b3c404
Queue commitmentsha256:ae0cb263464d5d0a7ae46277474ed3580302c825c4830110a70d1685cd88d2b8
Test lineagev2 · curated:prism-garden:v2
Evidence recipev1 · prism-garden-standard-demonstration
Review queue9/9 locked
Record hashsha256:ac15e88aba67131e7455e2763254ef979010e97a6a7f67ff1c5438d13c53043d
The frozen assignmentShow exact prompt

Build a one-screen interactive web experience called PRISM GARDEN. Use a dark near-black background. Users click or drag to plant luminous geometric flowers; each flower should bloom with motion and slight variation. Include a small visible bloom counter and a Reset control. Make it polished, responsive, and immediately understandable. Use only HTML, CSS, and vanilla JavaScript in a single file. No external assets, fonts, or libraries. It must work offline. Spend time on the experience, not explanations.

Limits of the claimRead all caveats
  • Five generations are enough to expose material variability, not to estimate a stable population distribution.
  • Two duplicate reviews measure consistency inside this review session only.
  • A complete HTML submission can still fail every requirement; submission completion and requirement compliance are reported separately.