More reasoning did not mean better work.

Across the valid common tasks, medium matched every higher reasoning setting while taking less time and fewer input tokens.

294 measured runs across GPT-5.6 Luna, Terra, and Sol through ChatGPT subscription access in Codex.

Published 01 Aug 2026

How the evidence changed

  1. 294measured rows retained in the historical ledger
  2. 50rows excluded after two defective checks were found
  3. 210rows eligible for the corrected common policy view

The decision

Start ordinary bounded work at medium. Escalate only when the task names a hard problem, costly failure, or unresolved effort decision.

42/42
accepted at medium, high, xhigh, and max
No winner
the valid tasks did not separate Luna, Terra, and Sol
Medium first
verify the work before paying for more reasoning

The result stayed flat. The cost did not.

Follow the same accepted result through each reasoning level. The bars show mean time, while the figures preserve the measured token cost.

Accepted work stays level while time and input tokens rise.

Medium42/4235.61s / 54,478 tokens
High42/4237.54s / 58,591 tokens
Extra High42/4247.17s / 66,298 tokens
Max42/4257.11s / 72,143 tokens
  1. Medium reached the measured ceiling

    Medium accepted every valid common task. This is the first complete result in the comparison.

  2. High added cost, not accepted work

    High also accepted 42 of 42, but required 37.54 seconds and 58,591 input tokens per accepted result.

  3. Extra High added cost, not accepted work

    Extra High also accepted 42 of 42, but required 47.17 seconds and 66,298 input tokens per accepted result.

  4. Max added cost, not accepted work

    Max also accepted 42 of 42, but required 57.11 seconds and 72,143 input tokens per accepted result.

A confident score measured the wrong thing.

The page could remain correct while the old checker changed its verdict. This is why the historical frontend rows cannot select a model.

Clear messageReadable supporting copy
Visible evidenceResponsive comparison
Observed layoutPassThe layout and implementation happened to match the old check.
  1. The page rendered as intended

    Desktop showed two readable regions and mobile stacked them. The user-facing outcome was correct.

  2. A harmless wrapper changed the score

    Adding a semantic container preserved the same layout, but the old checker inspected only direct children and reported failure.

  3. A repaired check reversed most failures

    The apparent aggregate moved from 7 of 33 under the old rule to 28 of 33 under a prototype repair.

  4. The historical rows were quarantined

    All 35 PUB8 rows remain in the ledger but no longer support routing. PUB8V2 is separately versioned and not yet measured on a model.