# grok 4.6 ranged from 8 to 92 across five Prism Garden runs

Five independent generations under one frozen condition produced a median craft score of 86; 4 met every requirement and 1 met none.

## Result

Generation-order scores: **86, 8, 92, 77, 89**. Median **86**; range **8–92**. 4/5 targets met every frozen requirement; 1/5 met none.

| Generation | Craft | Compliance | Submission outcome | Recipe evidence | Reported cost |
|---:|---:|---:|---|---|---:|
| 1 | 86 | 100% | completed submission | captured | $0.0624 |
| 2 | 8 | 0% | completed submission | failed | $0.0944 |
| 3 | 92 | 100% | completed submission | captured | $0.0950 |
| 4 | 77 | 100% | completed submission | captured | $0.0746 |
| 5 | 89 | 100% | completed submission | captured | $0.1111 |

## Method

Five independent generations used the same frozen prompt, model, route policy, T=0.7, 16,000 output-token ceiling, and Evaluation v3 scorecard. The blind queue contained 9 items: five targets, 2 salts, and 2 hidden duplicates. Identities and duplicate links remained hidden until all 9 reviews were locked.

The separate cold-rescore rehearsal scored 6 points higher than the original review. Duplicate-review deltas were 0 and 0. These describe reviewer consistency inside this review session only.

## Limitations

- Five generations are enough to expose material variability, not to estimate a stable population distribution.
- Two duplicate reviews measure consistency inside this review session only.
- A complete HTML submission can still fail every requirement; submission completion and requirement compliance are reported separately.

---

Study RP-MSROO2DW-E03530 · sha256:ac15e88aba67131e7455e2763254ef979010e97a6a7f67ff1c5438d13c53043d
