glm-5.3 and gemini-3.7-flash: an incomplete Orbit Radio run
One model produced a reviewable artifact and the other did not. Proofrun preserves the artifact assessment and the failure record without declaring an automatic head-to-head winner.
The result,
shown before explained.
“Hard to review this incomplete artifact, barely anything to interact with, from what I can see the UI is okay but uninspired.”Locked before model reveal
Requirements first.
Weighted quality second.
| Requirement | A | B |
|---|---|---|
| The central dial responds to turning or dragging. | Unscored | Missed |
| Five distinct fictional stations are available. | Unscored | Missed |
| Station name, frequency, color, and visualization change together. | Unscored | Missed |
| Previous and next controls work consistently. | Unscored | Missed |
| The current station is visually unmistakable. | Unscored | Missed |
| The single-file instrument remains responsive and works offline. | Unscored | Missed |
Visual design · 30%
Hierarchy, composition, typography, color, polish, and coherent visual judgment.
Interaction & usability · 30%
Discoverability, responsiveness, feedback, control quality, and interaction feel.
Technical robustness · 30%
Runtime correctness, edge-case handling, performance, responsiveness, and code reliability.
Originality · 10%
Useful creative choices that go beyond a generic first solution without harming clarity.
The claim, with its
conditions attached.
Immutable attempts and publication recordInspect provenance
This Results Brief is rendered from a sanitized, allowlisted record rather than AI-authored article prose. Credentials, raw responses, generated source, executable artifacts, and private reviewer notes remain in the private Studio.
A #1 original: invalid or non reviewable artifact (token ceiling, Z.AI) · B #1 original: truncated submission (token ceiling, Google AI Studio)