Contestant A · preserved capturedeepseek-v4-flash-0731 vs hy3: Weather Story
Turns the same small dataset into a useful, visually opinionated forecast.
Contestant A · preserved capture
Contestant B · preserved captureSelected and approved in Proofrun Studio. The interactive viewer renders the verified, inert scene payload; the preserved capture remains the canonical fallback.
The result
deepseek/deepseek-v4-flash-0731 · medium confidence · 92–90
Contestant A's result is visually much more interesting and polished.
What was tested
Turns the same small dataset into a useful, visually opinionated forecast.
Design a polished one-page weather experience for Helsinki using this fixed six-hour temperature sequence: 12 degrees, 13 degrees, 15 degrees, 14 degrees, 11 degrees, 9 degrees. Show the progression visually, highlight the warmest hour, and include wind and rain probability as invented but clearly labeled demo data. Add one meaningful hover or tap interaction. Use one offline HTML file with vanilla CSS and JavaScript; no libraries, APIs, or external assets.
Evidence to discuss
- Compare the two fixed screenshots and call out the most important visible difference.
- Explain requirement failures, runtime flags, controlled-probe findings, and attempt integrity.
- Distinguish this single run from a general model ranking.
Limitations
This is one blind comparison under Proofrun Protocol v0.7. It is evidence about this run, not a universal ranking.
Run PR-MSP08AE0-1E · sha256:828ab0d2a0f5ec64fa258964148301b79f9f4ce767da46382d0ce185e4216538
The claim, with its
conditions attached.
Method and publication recordInspect provenance
This page uses a sanitized, allowlisted record. Credentials, raw responses, generated source, executable artifacts, and private reviewer notes remain in the private Studio.
A #1 original: truncated submission (token ceiling, Baidu) · A #2 retry of #1: completed submission (normal, Decart) · B #1 original: completed submission (normal, Tencent)
Evidence analysis used openai/gpt-5.6-sol.