Model Release Profile 015 / Three frozen Solo runs / Published August 29, 2026

Hy4 preview completed two of three Proofrun tests—and the third stopped at the provider boundary

Across Proofrun’s first Model Release Profile, Tencent’s exact FP8 route scored 65 on Triage Desk and 61 on Verdant Terminal; Weather Story produced no artifact after an HTTP 429, so it remains incident evidence—not a model failure.

Bounded result2 completed artifacts and 1 provider incident.The slot scores remain separate. Proofrun does not publish a composite model score.
Modeltencent/hy4-previewExact routetencent/fp8 / fp8Protocolv0.8 revision 3ProfileMRP-001
Independent slot outcomes65 · INCIDENT · 61
Artifacts2/3
Known cost$0.1077
CompositeForbidden
01 / Three outcomes

One model route.
Three separate records.

The software anchor, Quick rotation, and Voxel rotation were selected before generation. Each retains its own score or terminal incident; none is averaged into a model-wide number.

Slot 01 / standard anchor65/100
Triage Desk: Matched initial state.
Fresh artifact state captured at the fixed 1280 × 720 review viewport.
Triage Desk

A very average result further weighed down by the lack of responsiveness in the UI.

Craft
65
Compliance
90%
Cost
$0.0600
Completed artifactPR-MTD5FJQT-3F6B5B
Slot 02 / quick rotationNo score
Provider incident

HTTP 429 before any reviewable artifact or usage receipt. No model score was assigned.

Weather Story

No artifact existed to assess. The terminal provider record is preserved without turning infrastructure failure into model evidence.

Craft
Compliance
Cost
Unknown
No submissionPR-MTD6103R-932CEA
Slot 03 / differentiated spatial rotation61/100
Verdant Terminal: Matched initial state · model camera.
Fresh artifact state captured at the fixed 1280 × 720 viewport from the contestant-authored camera; the reviewer inspection view was restored after capture.
Verdant Terminal

Average to slightly below average result containing flaws like the rails missing, issues with the stairs to the railway bridge and overall lack of detail.

Craft
61
Compliance
80%
Cost
$0.0477
Completed artifactPR-MTD7GPX4-2EF8BE

Three tests, three outcomes

Proofrun’s first Model Release Profile did not produce one tidy number. It produced three different records.

Tencent Hy4 preview completed Triage Desk with a craft score of 65/100 and 90% requirement compliance. It completed the Voxel Forge test Verdant Terminal at 61/100 and 80% compliance. Between them, Weather Story ended almost immediately with an HTTP 429 from the provider route and no model output at all.

Those results belong beside one another, but they should not be averaged. The two scores describe two artifacts under two different test rubrics. The Weather outcome describes an unsuccessful provider request, not the quality of a missing artifact. Proofrun therefore publishes the profile as three independent Solo records with no composite model score.

That distinction is the point of the profile format. A new model release invites a broad question—what can it actually make?—without giving us permission to turn a small, deliberately varied sample into a universal rating.

Triage Desk: reliable state, weaker presentation

The Triage Desk artifact completed the core support workflow coherently. It preserved all four frozen incidents, combined search and filters correctly, kept queue and detail state synchronized, blocked resolution when the internal note was blank, and retained the note and updated counts after a successful resolution.

That functional coherence is visible in the rubric. State robustness received 10/10 and workflow reasoning 7/10. Nine of ten frozen requirements passed.

The weaker side was the interface around that workflow. Information design received 6/10, while interaction and accessibility received 5/10. In the human review, the layout felt less resolved than the underlying state model, and the narrow-width requirement failed. The result was a capable operational artifact whose behavior was stronger than its presentation.

The final Triage Desk scores were 65/100 for craft and 90% for compliance, with medium confidence. The request ended normally after about 5 minutes 6 seconds. It used 23,775 completion tokens, including 16,523 reported reasoning tokens, at a reported cost of $0.059970849.

Weather Story: an incident, not a zero

Weather Story never reached the model-artifact stage. The pinned route returned HTTP 429 roughly one second after dispatch. Proofrun received no visible submission, no usage receipt, and no reviewable evidence.

The slot therefore has no craft score, no compliance percentage, and no screenshot. It is not scored as zero. A zero would claim that Hy4 attempted the task and produced work of the lowest possible quality; this record supports no such claim.

The incident still matters. A model release is experienced through its available route, and a route that cannot accept a request affects whether the model can be used at that moment. But reliability evidence and capability evidence are different things. Here we can report one observed provider incident. We cannot infer a general failure rate, and we cannot use the incident to judge Hy4’s ability to design a weather interface.

Because no usage receipt was returned, any charge associated with this failed request is unknown rather than assumed to be zero.

Verdant Terminal: efficient construction, under-resolved geometry

Verdant Terminal showed a different balance. Hy4 produced a valid, renderable SceneScript scene with a readable station cutaway, a recognizable train, a glass-roof structure, integrated vegetation, warm-to-cool material contrast, and a camera that held the major layers together.

The construction was economical. Command economy received 9/10, and composition received 7/10. The scene achieved a recognizable architectural idea without wasting its bounded command language.

The missing and ambiguous geometry kept it from feeling finished. The required two parallel rail lines were not visibly established. The elevated bridge existed, but its relationship to both rail lines and its stairs was only partially convincing. The luminous clock or departure-board landmark was also judged partial, and the scene lacked some of the detail needed to make the terminal feel fully authored rather than broadly sketched.

Verdant Terminal finished at 61/100 craft and 80% compliance, again with medium confidence. The request ended normally after about 4 minutes 46 seconds. It used 18,779 completion tokens, including 17,193 reported reasoning tokens, at a reported cost of $0.047725219.

What the profile tells us—and what it does not

Across the two completed runs, Hy4 preview showed that it could sustain long reasoning and return valid artifacts before the frozen completion ceilings. Neither successful request was truncated. In Triage Desk, its strongest quality was coherent application state. In Verdant Terminal, it was compact structural construction. In both, the main limitations appeared in the final layer of design resolution: responsive presentation in one case, spatial specificity and detail in the other.

The profile also preserves a less flattering operational fact: one of the three selected requests never produced an artifact. We report that outcome because hiding it would make the release look cleaner than the observed record. We do not call it a model-quality failure because the evidence does not justify that interpretation.

The known reported cost for the two completed artifacts was $0.107696068. The Weather charge is unknown. All three slots used the same pinned tencent/fp8 route, High native reasoning, no fallbacks, and one initial paid attempt. Proofrun retained its frozen temperature of 0.7 even though Tencent recommended 0.9, so the difference is disclosed rather than adjusted after seeing the results.

This is still a small profile: two completed artifacts and one provider incident. It does not establish a stable quality distribution, a reliability rate, or a general ranking of Hy4 against other models. It does give us three honest run-level outcomes from a prospectively selected software anchor, Quick rotation, and spatial rotation—and a format capable of keeping those outcomes together without pretending they are the same measurement.

Independent scorecards

Scores stay attached
to their own tests.

Triage Desk65/100 craft · 90% compliance
Frozen requirementJudgment
Exactly the four frozen incidents and their supplied facts are present, with every incident initially Open.Met
The queue makes urgency, assignment, status, and useful aggregate counts easy to scan.Met
Text search, status filters, and severity filtering compose correctly and include a clear empty-result state.Met
Selecting an incident keeps the queue context visible and reveals its facts, customer report, and activity timeline.Met
Severity and owner changes remain synchronized between the selected detail, queue, filters, and counts.Met
Attempting to resolve with a blank internal note is blocked and produces visible, actionable guidance.Met
A successful resolution preserves the note in the timeline and clearly changes the incident to Resolved.Met
After resolution, the selected detail, queue row, filters, and aggregate counts remain mutually consistent.Met
The workflow remains usable at narrow widths and provides semantic labels, keyboard operation, and visible focus.Missed
The artifact is one responsive, self-contained HTML document that works offline without external dependencies.Met

Workflow reasoning · 30%

Whether selection, assignment, notes, validation, resolution, and dependent state changes form one coherent operational flow.

Score / 107

Information design · 25%

Scanability, hierarchy, density, prioritization, and the relationship between the queue and incident detail.

Score / 106

Interaction & accessibility · 25%

Control clarity, keyboard use, focus visibility, responsive adaptation, validation, and useful feedback.

Score / 105

State robustness · 20%

Consistency of incident data, filters, counts, selected detail, and edge cases as the workflow changes.

Score / 1010
Weather StoryProvider incident / no score
No artifact judgmentProvider returned error · HTTP 429 · provider incident or capacity. This remains infrastructure evidence, not a zero.
Verdant Terminal61/100 craft · 80% compliance
Frozen requirementJudgment
A grounded station hall has a clearly readable cutaway interior rather than only an exterior facade.Met
Two parallel rail lines visibly pass through the station hall.Missed
A stationary train has a recognizable front or locomotive and at least two visibly linked cars.Met
An elevated pedestrian bridge spans both rail lines and visibly reaches the platform areas at both ends by stairs or ramps.Partial
A repeating structural frame supports a visibly glass roof without obscuring the main interior.Met
Botanical details are integrated across multiple distinct parts of the station rather than confined to one token planter.Met
A luminous clock or departure board forms a strong interior landmark.Partial
Warm inhabited lighting contrasts deliberately with cooler rail, frame, or glass materials.Met
The selected camera keeps the tracks, train, bridge, roof, vegetation, and vertical layering readable together.Met
The response produces a renderable SceneScript scene within the frozen validator limits.Met

Spatial reasoning · 30%

Coherent three-dimensional placement, scale, connection, support, and navigable relationships.

Score / 106

Composition · 30%

Silhouette, depth, focal hierarchy, camera choice, balance, and use of negative space.

Score / 107

Scene craft · 25%

Material contrast, detail distribution, repetition, variation, and voxel-specific visual character.

Score / 105

Command economy · 15%

Complexity achieved through valid, bounded SceneScript rather than wasteful repetition or ignored operations.

Score / 109
Verdant Terminal: Deterministic context view.
Deterministic context viewSupplementary 1280 × 720 perspective context view. Proofrun follows the standardized camera direction, retargets toward the scene bounds, uses a moderate lens, and deliberately favors readable subject scale over complete edge coverage. The authored environment and canonical camera evidence remain unchanged.
Verdant Terminal: Deterministic orthographic survey.
Deterministic orthographic surveySupplementary 1280 × 720 orthographic survey fitted to the complete voxel bounds from the standardized camera direction. It preserves the authored environment while removing perspective shrinkage; the canonical camera remains standardized primary evidence.
Method / Profile record

The selection and
failures stay visible.

Triage Desk5:06$0.0600 reported cost
Verdant Terminal4:46$0.0477 reported cost
Weather StoryHTTP 429provider incident · charge unknown
Profile and publication recordInspect provenance

The journal renders one allowlisted profile record, its human-approved article, and selected static evidence. It executes no generated contestant HTML. Raw responses, reasoning text, credentials, router metadata, generation identifiers, and private reviewer notes remain private.

ProfileMRP-001
AuthorizationMLA-7828660B54A61C38
Suitev1 revision 2
Protocol hashsha256:6bda63bb3ec8881dc9d84ffdc1bb1e4fbfd3faee9d850680bfaf862f20a5e1d6
Profile hashsha256:8d271fb4180236a10c29a826cb3ec16f88c62871c2cb443d6f5c522219b7235b
Publication hashsha256:41385fcdec0cbde7fc0815c2b4d1693eaa0af03eb920e876ea1f2dd1e974a9f9
Sampling disclosure

High reasoning, T=0.7, pinned tencent/fp8, no fallbacks, one initial attempt per slot. Tencent recommended T=0.9; Proofrun retained its frozen suite condition.

AI assistance

Article drafting assistance used openai/gpt-5.6-sol; the exported revision was explicitly approved by the human editor.

Limits of the claimRead all caveats
  • Two completed artifacts and one provider incident cannot establish a stable distribution of model quality or reliability.
  • The Weather Story slot produced no submission and therefore supports no claim about Hy4's ability on that test.
  • The profile used Proofrun's frozen temperature of 0.7 rather than the vendor-recommended 0.9; the difference is disclosed, not corrected after seeing results.
  • Known reported cost excludes any unreceipted charge that may have resulted from the Weather provider incident.