One ceiling was doing four jobs.
Across 16 completed one-file HTML attempts, the median visible artifact used approximately 6,804 tokens and the largest complete artifact used 9,387. Three preserved partial artifacts reached approximately 7,567, 11,882, 12,385 visible tokens before a length stop.
The immediate correction followed two Qwen3.8 Flash Triage Desk conditions under revision 1. High reasoning used 12,118 reasoning tokens and approximately 11,882 visible tokens inside one shared 24,000-token ceiling. The artifact was reviewable but truncated. Medium reasoning used all 24,000 reported completion tokens as reasoning and returned no visible artifact.
Artifact first. Reasoning added.
The test scope now defines the visible-artifact budget. Proofrun adds the selected reasoning allowance to produce the total completion ceiling.
| Test scope | Artifact budget |
|---|---|
| quick | 16,000 tokens |
| standard | 24,000 tokens |
| deep | 32,000 tokens |
| extended | 48,000 tokens |
| Reasoning | Allowance |
|---|---|
| none | 0 tokens |
| low | 4,000 tokens |
| medium | 8,000 tokens |
| high | 16,000 tokens |
| very high | 24,000 tokens |
A Standard test at High therefore receives 24,000 visible-artifact tokens plus a 16,000-token reasoning allowance: 40,000 total. The largest ordinary combination is 72,000 tokens.
Protected or target—never implied.
The reasoning cap is enforced and the visible-artifact budget is protected.
The visible-artifact amount is disclosed as a target, not a guaranteed reserve.
A blind Duel cannot mix those two mechanisms under one shared condition.
Earlier evidence stays earlier evidence.
Revision-1 shared-ceiling results and the unused revision-2 50/50 contract remain unchanged and are not reinterpreted under revision 3.
Revision 3 is prospective. It creates new conditions; it does not repair or relabel an earlier attempt.
A condition, not a promised story.
The first proposed application is qwen/qwen3.8-flash on Triage Desk v1, pinned to the Alibaba route: 24,000 artifact tokens plus 16,000 High-reasoning tokens, or 40,000 total.
It remains one possible attempt. This note authorizes neither its launch nor its publication.
More room is not a quality guarantee.
- A larger budget reduces avoidable truncation; it cannot guarantee a complete or high-quality artifact.
- Provider-native effort does not expose an enforceable reasoning-token maximum.
- Tokenization varies across models, so equal token counts do not imply equal character counts.
- Exhausting a generous declared artifact budget remains valid evidence about scope discipline under that condition.
Inspect the change.
Protocol Note content hashsha256:82f0298b980b5a8ddc81679d0ed953bc840c2062ce849220b3c4cb8f65a848bc