Field test 004 / Voxel Forge / Published August 12, 2026

Fable’s Sky Archive scored 74: a strong beacon undermined by an extra island and unfinished bridges

The solo Voxel Forge run produced valid, economical SceneScript with a clear central landmark, but the locked evaluation penalized its island count, bridge connections, dark palette, and uneven composition.

Bounded resultanthropic/claude-fable-5 produced a 74/100 artifact in this run.One run. One frozen brief. No broader model rating is implied.
5 min readAugust 12, 2026Voxel Forgeproofrun-v0.7
Modelanthropic/claude-fable-5
Confidencemedium
Artifact score74/100
Requirements met5/7
Publication primary

Selected and approved in Proofrun Studio. The interactive viewer renders the verified, inert scene payload; the preserved capture remains the canonical fallback.

Result: Anthropic’s Claude Fable 5 earned a locked human artifact score of 74, with medium confidence, in this solo Sky Archive test. This is one frozen artifact evaluation, not a universal model rating or a comparison with another model.

The scene’s clearest success is its luminous central archive: a tall stack of bright white, cyan, and gold elements, capped by a beam, that dominates the surrounding voxel landscape. The submission was renderable, used reusable SceneScript structures effectively, and included botanical and sculptural detail across its floating landmasses.

What held it back was not a lack of ideas. It was the control of those ideas. The locked evaluator found that the deep-purple background made the scene too dark and that the bridges from two adjacent islands did not properly meet the central island. The evaluator also failed the exact island-count requirement.

What the model made

The prompt asked for a dramatic floating “Sky Archive” with exactly three distinguishable islands, a taller central beacon, at least two deliberate routes, visible secondary detail, material contrast, depth, and economical construction.

Fable produced a central archive platform surrounded by three satellite landmasses. The editor’s notebook identifies four discrete island groups in total: the central archive island plus three peripheral islands. The satellites were differentiated by features including flowers, an obelisk, and trees.

That arrangement created an interpretation problem. The editor noted that the model may have read “three surrounding islands plus a central archive” as permission to place the archive on a fourth island. The frozen requirement, however, was judged as demanding exactly three total, and the locked human evaluation marked it fail. That judgment remains authoritative.

The central landmark fared much better. It was the tallest and strongest focal point, satisfying the beacon requirement. The scene also passed requirements for botanical or sculptural details, material contrast, camera depth and readability, and production of a renderable SceneScript artifact.

Two route types were visibly authored: a segmented plank bridge and a line of luminous stepping stones. The technical audit found repeated plank instances and four stepping-stone instances in the source. But visible intent did not fully translate into spatial completion. The locked evaluator judged the connection requirement partial because the routes did not properly connect to the central island.

Score profile

The locked overall score was 74. Its criterion profile shows where the scene was strongest and where it lost control:

CriterionWeightScore
Instruction fit20%9/10
Spatial reasoning25%8/10
Composition25%6/10
Scene craft20%6/10
Command economy10%9/10

The high instruction-fit and command-economy scores reflect a scene that contained most requested elements and constructed them through bounded, reusable SceneScript. The lower composition and scene-craft scores align with the darkness, imperfect bridge joins, and difficulty presenting all of the scene’s elements cleanly at once.

What the technical audit added

The accepted AI findings support several parts of the human assessment without replacing it.

The source defined reusable groups for islands, vegetation, bridge planks, stepping stones, lanterns, pillars, and sculptural objects. It also used an array for a patterned flower patch. That is concrete evidence for the 9/10 command-economy score, although this run does not establish how efficient the source would be relative to other possible implementations.

The audit also confirmed that the three peripheral landmasses are distinct in wider inspection views and that the beacon remains visually dominant. Trees, bushes, flowers, hanging greenery, lantern-like accents, crystals, grass caps, and luminous rings provide secondary detail beyond the main architecture.

The uncomfortable evidence concerns presentation. Dark rails, island undersides, pillars, and parts of the archive rendered close to black against the violet background, compressing structural detail in shadow. This supports the locked evaluator’s criticism of the scene’s darkness.

There is also a camera caveat. The locked human evaluation passed the camera requirement, but the accepted audit warning found that the contestant-authored capture cut off the top of the beacon shaft and much of the nearest island. In that capture, only two peripheral landmasses were readily distinguishable; the wider inspection views made the full arrangement clearer. The audit warning does not alter the locked pass, but it shows that the scene was easier to understand during inspection than through its authored framing alone.

Time, cost, and run integrity

The completed generation took 114.5 seconds of wall time. Recorded usage was 834 input tokens, 8,567 completion tokens, and 1,322 reasoning tokens, at a recorded cost of $0.43669 for the completed generation.

A preceding attempt produced no submission because OpenRouter returned an HTTP 404 routing rejection. The protocol record classifies this as a harness catalog-parameter compatibility incident rather than a contestant submission failure. The linked retry completed normally through Amazon Bedrock, with no recorded runtime errors or artifact issues.

The editor’s notebook records surprise that the first model generation to produce an artifact yielded valid SceneScript. That technical validity mattered: the model did not need to be rescued from a malformed final scene. Its shortcomings were visible compositional decisions inside an otherwise working artifact.

The editor’s view

The editor liked most of the constructed scene but disliked the deep-purple background and the imperfect bridge connections. The central beacon, varied island identities, and use of two different route types gave the artifact character. At the same time, the low ambient lighting made close inspection harder than it needed to be.

That tension defines this run. Fable built a scene with a strong concept, clear hierarchy, varied detail, and disciplined source reuse. It also allowed an extra island, incomplete connections, and shadow-heavy presentation to weaken the final read. The result is good, as the locked rationale says, but not fully resolved.

Limitations

This report covers one frozen solo artifact and should not be generalized into a broad rating of the model. No runtime probe was supplied, so camera controls, navigation, collision behavior, hidden geometry, and uncaptured viewpoints were not tested. Renderability is supported by the completed run and the absence of recorded runtime errors, but validator semantics and material defaults were not independently inspected.

The AI evidence analysis also had weak calibration measurements: repeat consistency was 0.111, swapped-packet consistency was 0.111, and position sensitivity was 0.889. Its accepted findings are therefore used only as supporting evidence; the locked human score, confidence, requirement judgments, and rationale remain authoritative.

Assistance note: Evidence analysis was provided by openai/gpt-5.6-sol. Article drafting assistance was provided by openai/gpt-5.6-sol.

Method / Publication record

The claim, with its
conditions attached.

Model artifact1:54$0.4367 reported cost
Artifact score74/1005/7 requirements met
Protocolv0.7curated:voxel-forge-sky-archive:v1
Method and publication recordInspect provenance

This page uses a sanitized, allowlisted record. Credentials, raw responses, generated source, executable artifacts, and private reviewer notes remain in the private Studio.

Run recordPR-MSPLJNSY-94
Protocolproofrun-v0.7
Test lineagev1 · curated:voxel-forge-sky-archive:v1
Identityrevealed
Evaluationlocked
Outcomesolo artifact
Record hashsha256:ec293d78d56e019fb1fab5028d35a103c981164b66f563ac7666cf6257bae598
Immutable attempt ledger

A #1 original: no submission (model or provider error) · A #2 retry of #1: completed submission (normal, Amazon Bedrock)

Protocol exceptions

catalog-parameter-compatibility: The original attempt requested temperature from a model whose frozen OpenRouter catalog capabilities did not support it. The routing rejection is a harness compatibility incident, not a contestant submission failure. · voxel-camera-provenance-upgrade: An earlier native voxel capture predates frozen camera provenance and may reflect the reviewer inspection orbit. It remains preserved; standardized evidence was recaptured from the contestant-authored camera or deterministic auto-fit under voxel-renderer-v1.

AI assistance

Evidence analysis used openai/gpt-5.6-sol. Article drafting used openai/gpt-5.6-sol; the final revision was human approved.

Download the hashed publication record