PN-001Published August 27, 2026Methodology · not a review Issue

Separating reasoning and artifact budgets

Proofrun now assigns a visible-artifact budget from the test scope and adds a separate reasoning allowance, so hidden deliberation does not silently consume the artifact's declared room.

Published prospectivelyNo public revision-3 result existed when this note was published.
Protocolproofrun-review
Version0.8
Revision3
Effective2026-08-27
Manifest hashsha256:6bda63bb3ec8881dc9d84ffdc1bb1e4fbfd3faee9d850680bfaf862f20a5e1d6

One ceiling was doing four jobs.

Across 16 completed one-file HTML attempts, the median visible artifact used approximately 6,804 tokens and the largest complete artifact used 9,387. Three preserved partial artifacts reached approximately 7,567, 11,882, 12,385 visible tokens before a length stop.

The immediate correction followed two Qwen3.8 Flash Triage Desk conditions under revision 1. High reasoning used 12,118 reasoning tokens and approximately 11,882 visible tokens inside one shared 24,000-token ceiling. The artifact was reviewable but truncated. Medium reasoning used all 24,000 reported completion tokens as reasoning and returned no visible artifact.

Reasoning allowance, artifact room, total ceiling, and cost control should not be one ambiguous number.

Artifact first. Reasoning added.

The test scope now defines the visible-artifact budget. Proofrun adds the selected reasoning allowance to produce the total completion ceiling.

Test scopeArtifact budget
quick16,000 tokens
standard24,000 tokens
deep32,000 tokens
extended48,000 tokens
ReasoningAllowance
none0 tokens
low4,000 tokens
medium8,000 tokens
high16,000 tokens
very high24,000 tokens

A Standard test at High therefore receives 24,000 visible-artifact tokens plus a 16,000-token reasoning allowance: 40,000 total. The largest ordinary combination is 72,000 tokens.

Protected or target—never implied.

Direct reasoning budget

The reasoning cap is enforced and the visible-artifact budget is protected.

Provider-native effort

The visible-artifact amount is disclosed as a target, not a guaranteed reserve.

A blind Duel cannot mix those two mechanisms under one shared condition.

Earlier evidence stays earlier evidence.

Revision-1 shared-ceiling results and the unused revision-2 50/50 contract remain unchanged and are not reinterpreted under revision 3.

Revision 3 is prospective. It creates new conditions; it does not repair or relabel an earlier attempt.

A condition, not a promised story.

The first proposed application is qwen/qwen3.8-flash on Triage Desk v1, pinned to the Alibaba route: 24,000 artifact tokens plus 16,000 High-reasoning tokens, or 40,000 total.

It remains one possible attempt. This note authorizes neither its launch nor its publication.

More room is not a quality guarantee.

  • A larger budget reduces avoidable truncation; it cannot guarantee a complete or high-quality artifact.
  • Provider-native effort does not expose an enforceable reasoning-token maximum.
  • Tokenization varies across models, so equal token counts do not imply equal character counts.
  • Exhausting a generous declared artifact budget remains valid evidence about scope discipline under that condition.

Inspect the change.

Protocol Note content hashsha256:82f0298b980b5a8ddc81679d0ed953bc840c2062ce849220b3c4cb8f65a848bc