Field test 001 / Physics toy / Published August 9, 2026

Terra’s steadier physics beat Sonnet 5’s audio edge in Gravity Desk, 87–73

In one blind, frozen comparison, contestant B won with high confidence because stronger physics, control, and presentation outweighed A’s more responsive-feeling sound feedback.

The short versionTerra won on feel, control, and finish. Sonnet’s clearest edge was sound.One run. One frozen brief. No broader ranking is implied.
9 min readAugust 9, 2026Physics toyproofrun-v0.6
Verdictopenai/gpt-5.6-terra
Confidencehigh
Blind score73:87
Margin+14
01 / The verdict

Same brief.
Different finish.

Both models delivered the complete one-screen physics toy: clickable marbles, gravity, collisions, working controls, and a live object count. Completion did not decide this comparison.

Terra’s movement felt steadier and its controls and presentation felt more resolved. Sonnet added sound that made individual actions feel more responsive, but that advantage did not outweigh the gap in physics and finish.

Why Terra wonMore stable motion and stronger overall control.
Sonnet’s edgeSound gave clicks and collisions more feedback.
Important caveatTerra has no preserved automated probe for this run.
B's result is overall better because of the stabler physics, more control over the balls and better visuals overall.Evaluation locked before model reveal

Publication-primary captures show reviewer-created populated states. A contains 44 marbles and B contains 31, so the images compare presentation—not capacity.

02 / The scorecard

Both met the brief.
The quality was not equal.

Every requirement was judged before model identities were revealed. The weighted score then separated basic completion from how convincing the result felt in use.

Frozen requirements5 met / 1 partial each
RequirementAB
Clicking spawns a colorful glass-like marble.MetMet
Marbles fall and bounce within the viewport.MetMet
Marbles collide convincingly enough for the intended toy.MetMet
Gravity and Clear controls work during active simulation.MetMet
The live object count remains accurate.MetMet
The single-file canvas experience remains smooth, responsive, and offline.PartialPartial
Weighted criteriaA 73 / B 87

Requirement fit · 20%

How completely and accurately the artifact satisfies the frozen brief.

A / 1010
B / 1010

Physics & feel · 25%

Believability, responsiveness, collision behavior, and the satisfaction of motion.

A / 107
B / 1010

Technical robustness · 25%

Runtime correctness, edge-case handling, performance, responsiveness, and code reliability.

A / 106
B / 107

Visual design · 20%

Hierarchy, composition, typography, color, polish, and coherent visual judgment.

A / 106
B / 107

Interaction & usability · 10%

Discoverability, responsiveness, feedback, control quality, and interaction feel.

A / 108
B / 1010
What the code suggested

Both models implemented substantive collision systems. Terra used a fixed 1/120-second accumulator, three collision passes, low-speed restitution, and tangential friction. Those choices help explain the result, but the runtime judgment—not code aesthetics—determined the score.

03 / What held up

The winner was stronger.
It was not spotless.

Terra’s populated capture shows a fixed footer crossing the settled marbles. It also lacks a recorded automated interaction probe. Its controls passed human evaluation, but the packet cannot independently confirm those paths.

Sonnet’s probe successfully exercised a click, drag, five-point burst, and Clear at a narrow viewport without runtime errors. That confirms the input paths executed; it does not prove visual correctness or performance near the 260-marble limit.

The limitNo motion telemetry or video was preserved. “Steadier physics” remains a bounded human judgment from this run, not a measured simulation claim.
Contestant Apassed
Narrow viewport

536 x 520 viewport exercised

Single click

Dispatched one complete pointer and mouse click sequence

Short drag

Dispatched a bounded six-step diagonal drag

Five-point burst

Dispatched five bounded clicks across the interaction surface

Reset control

Activated 🧹 Clear

Contestant Bnot recorded

No automated interaction probe was preserved for this contestant.

A / Performance risk

Pairwise collision checks could become expensive near the 260-marble cap. The run did not establish an observed slowdown.

A / Accessibility

The canvas has no keyboard spawning route or accessible name, and viewport scaling is disabled.

B / Visual overlap

The fixed footer competes with the settled marble row during ordinary play.

04 / The run record

The claim, with its
conditions attached.

Sonnet 51:46$0.1229 reported cost
Terra1:32$0.1361 reported cost
Difference14.6sTerra finished sooner
Limits of the claimRead the caveats

This was one blind comparison under frozen conditions, not a universal ranking. Static captures cannot establish bounce consistency, tunneling, pile stability, or collision satisfaction. Terra lacks a recorded automated probe; Sonnet’s probe was too limited to establish high-population performance. Audio, physical-device touch behavior, reduced-motion preferences, and low-power hardware were not systematically tested.

The frozen assignmentShow the exact prompt

Build a playful one-screen physics toy called GRAVITY DESK. Clicking spawns colorful glass-like marbles that fall, bounce off the viewport edges, and collide convincingly enough to feel satisfying. Include Gravity and Clear controls plus a live object count. Use a single offline HTML file with Canvas and vanilla JavaScript, no libraries or assets. Prioritize smoothness, polish, and immediate playability.

Method and publication recordInspect provenance

The public page uses a sanitized, allowlisted record. Credentials, raw responses, generated source, executable artifacts, and private reviewer notes remain in the private Studio.

Run recordPR-MSKKQ2K6-34
Protocolproofrun-v0.6
Test lineagev1 · curated:gravity-desk:v1
Identityrevealed
Evaluationlocked
Record hashsha256:ceb35e3680ce91601e1004a220a487ef6beea0fdb36e5448b8ef47e6f5a047e3
AI assistance

Evidence analysis used openai/gpt-5.6-sol while identities were hidden; 6 findings were accepted and 3 rejected. The article draft used openai/gpt-5.6-sol and was approved by the human editor. Recorded assistance cost: $0.4538.

Download the hashed publication record