
A very average result further weighed down by the lack of responsiveness in the UI.
- Craft
- 65
- Compliance
- 90%
- Cost
- $0.0600
Across Proofrun’s first Model Release Profile, Tencent’s exact FP8 route scored 65 on Triage Desk and 61 on Verdant Terminal; Weather Story produced no artifact after an HTTP 429, so it remains incident evidence—not a model failure.
The software anchor, Quick rotation, and Voxel rotation were selected before generation. Each retains its own score or terminal incident; none is averaged into a model-wide number.

A very average result further weighed down by the lack of responsiveness in the UI.
HTTP 429 before any reviewable artifact or usage receipt. No model score was assigned.
No artifact existed to assess. The terminal provider record is preserved without turning infrastructure failure into model evidence.

Average to slightly below average result containing flaws like the rails missing, issues with the stairs to the railway bridge and overall lack of detail.
Proofrun’s first Model Release Profile did not produce one tidy number. It produced three different records.
Tencent Hy4 preview completed Triage Desk with a craft score of 65/100 and 90% requirement compliance. It completed the Voxel Forge test Verdant Terminal at 61/100 and 80% compliance. Between them, Weather Story ended almost immediately with an HTTP 429 from the provider route and no model output at all.
Those results belong beside one another, but they should not be averaged. The two scores describe two artifacts under two different test rubrics. The Weather outcome describes an unsuccessful provider request, not the quality of a missing artifact. Proofrun therefore publishes the profile as three independent Solo records with no composite model score.
That distinction is the point of the profile format. A new model release invites a broad question—what can it actually make?—without giving us permission to turn a small, deliberately varied sample into a universal rating.
The Triage Desk artifact completed the core support workflow coherently. It preserved all four frozen incidents, combined search and filters correctly, kept queue and detail state synchronized, blocked resolution when the internal note was blank, and retained the note and updated counts after a successful resolution.
That functional coherence is visible in the rubric. State robustness received 10/10 and workflow reasoning 7/10. Nine of ten frozen requirements passed.
The weaker side was the interface around that workflow. Information design received 6/10, while interaction and accessibility received 5/10. In the human review, the layout felt less resolved than the underlying state model, and the narrow-width requirement failed. The result was a capable operational artifact whose behavior was stronger than its presentation.
The final Triage Desk scores were 65/100 for craft and 90% for compliance, with medium confidence. The request ended normally after about 5 minutes 6 seconds. It used 23,775 completion tokens, including 16,523 reported reasoning tokens, at a reported cost of $0.059970849.
Weather Story never reached the model-artifact stage. The pinned route returned HTTP 429 roughly one second after dispatch. Proofrun received no visible submission, no usage receipt, and no reviewable evidence.
The slot therefore has no craft score, no compliance percentage, and no screenshot. It is not scored as zero. A zero would claim that Hy4 attempted the task and produced work of the lowest possible quality; this record supports no such claim.
The incident still matters. A model release is experienced through its available route, and a route that cannot accept a request affects whether the model can be used at that moment. But reliability evidence and capability evidence are different things. Here we can report one observed provider incident. We cannot infer a general failure rate, and we cannot use the incident to judge Hy4’s ability to design a weather interface.
Because no usage receipt was returned, any charge associated with this failed request is unknown rather than assumed to be zero.
Verdant Terminal showed a different balance. Hy4 produced a valid, renderable SceneScript scene with a readable station cutaway, a recognizable train, a glass-roof structure, integrated vegetation, warm-to-cool material contrast, and a camera that held the major layers together.
The construction was economical. Command economy received 9/10, and composition received 7/10. The scene achieved a recognizable architectural idea without wasting its bounded command language.
The missing and ambiguous geometry kept it from feeling finished. The required two parallel rail lines were not visibly established. The elevated bridge existed, but its relationship to both rail lines and its stairs was only partially convincing. The luminous clock or departure-board landmark was also judged partial, and the scene lacked some of the detail needed to make the terminal feel fully authored rather than broadly sketched.
Verdant Terminal finished at 61/100 craft and 80% compliance, again with medium confidence. The request ended normally after about 4 minutes 46 seconds. It used 18,779 completion tokens, including 17,193 reported reasoning tokens, at a reported cost of $0.047725219.
Across the two completed runs, Hy4 preview showed that it could sustain long reasoning and return valid artifacts before the frozen completion ceilings. Neither successful request was truncated. In Triage Desk, its strongest quality was coherent application state. In Verdant Terminal, it was compact structural construction. In both, the main limitations appeared in the final layer of design resolution: responsive presentation in one case, spatial specificity and detail in the other.
The profile also preserves a less flattering operational fact: one of the three selected requests never produced an artifact. We report that outcome because hiding it would make the release look cleaner than the observed record. We do not call it a model-quality failure because the evidence does not justify that interpretation.
The known reported cost for the two completed artifacts was $0.107696068. The Weather charge is unknown. All three slots used the same pinned tencent/fp8 route, High native reasoning, no fallbacks, and one initial paid attempt. Proofrun retained its frozen temperature of 0.7 even though Tencent recommended 0.9, so the difference is disclosed rather than adjusted after seeing the results.
This is still a small profile: two completed artifacts and one provider incident. It does not establish a stable quality distribution, a reliability rate, or a general ranking of Hy4 against other models. It does give us three honest run-level outcomes from a prospectively selected software anchor, Quick rotation, and spatial rotation—and a format capable of keeping those outcomes together without pretending they are the same measurement.
| Frozen requirement | Judgment |
|---|---|
| Exactly the four frozen incidents and their supplied facts are present, with every incident initially Open. | Met |
| The queue makes urgency, assignment, status, and useful aggregate counts easy to scan. | Met |
| Text search, status filters, and severity filtering compose correctly and include a clear empty-result state. | Met |
| Selecting an incident keeps the queue context visible and reveals its facts, customer report, and activity timeline. | Met |
| Severity and owner changes remain synchronized between the selected detail, queue, filters, and counts. | Met |
| Attempting to resolve with a blank internal note is blocked and produces visible, actionable guidance. | Met |
| A successful resolution preserves the note in the timeline and clearly changes the incident to Resolved. | Met |
| After resolution, the selected detail, queue row, filters, and aggregate counts remain mutually consistent. | Met |
| The workflow remains usable at narrow widths and provides semantic labels, keyboard operation, and visible focus. | Missed |
| The artifact is one responsive, self-contained HTML document that works offline without external dependencies. | Met |
Whether selection, assignment, notes, validation, resolution, and dependent state changes form one coherent operational flow.
Scanability, hierarchy, density, prioritization, and the relationship between the queue and incident detail.
Control clarity, keyboard use, focus visibility, responsive adaptation, validation, and useful feedback.
Consistency of incident data, filters, counts, selected detail, and edge cases as the workflow changes.
| Frozen requirement | Judgment |
|---|---|
| A grounded station hall has a clearly readable cutaway interior rather than only an exterior facade. | Met |
| Two parallel rail lines visibly pass through the station hall. | Missed |
| A stationary train has a recognizable front or locomotive and at least two visibly linked cars. | Met |
| An elevated pedestrian bridge spans both rail lines and visibly reaches the platform areas at both ends by stairs or ramps. | Partial |
| A repeating structural frame supports a visibly glass roof without obscuring the main interior. | Met |
| Botanical details are integrated across multiple distinct parts of the station rather than confined to one token planter. | Met |
| A luminous clock or departure board forms a strong interior landmark. | Partial |
| Warm inhabited lighting contrasts deliberately with cooler rail, frame, or glass materials. | Met |
| The selected camera keeps the tracks, train, bridge, roof, vegetation, and vertical layering readable together. | Met |
| The response produces a renderable SceneScript scene within the frozen validator limits. | Met |
Coherent three-dimensional placement, scale, connection, support, and navigable relationships.
Silhouette, depth, focal hierarchy, camera choice, balance, and use of negative space.
Material contrast, detail distribution, repetition, variation, and voxel-specific visual character.
Complexity achieved through valid, bounded SceneScript rather than wasteful repetition or ignored operations.


The journal renders one allowlisted profile record, its human-approved article, and selected static evidence. It executes no generated contestant HTML. Raw responses, reasoning text, credentials, router metadata, generation identifiers, and private reviewer notes remain private.
High reasoning, T=0.7, pinned tencent/fp8, no fallbacks, one initial attempt per slot. Tencent recommended T=0.9; Proofrun retained its frozen suite condition.
Article drafting assistance used openai/gpt-5.6-sol; the exported revision was explicitly approved by the human editor.