More reasoning did not mean better work.

We ran 294 measured Codex tasks across GPT-5.6 Luna, Terra, and Sol. On the valid tasks, medium matched the higher reasoning settings while taking less time and fewer tokens. The practical starting point is medium, followed by verification. Escalate only when the risk or difficulty is clear.

Tested through ChatGPT subscription access in Codex. Common comparison: low, medium, high, xhigh, and max. Separate appendix: Terra and Sol ultra.

294 runs
common comparisons, targeted repeats, and separate ultra tests
42/42
valid common tasks passed at medium, high, xhigh, and max
No winner
the valid tasks did not separate Luna, Terra, and Sol

The decision: start ordinary bounded Codex work at medium. Use higher reasoning for a named hard problem or costly failure, not as a default. This is a finding about our July 2026 tasks and runtime, not a universal ranking of OpenAI models.

The operating policy

Codex presents reasoning as a dial. Turning that dial upward sounds safer: give the model more time, get better work. It also means longer waits and greater usage. Without evidence, a team can spend more on every task without knowing whether the result changed.

The benchmark was built to answer an everyday business question: which Codex setting should a team use before it knows whether a task is hard?

Our answer is a starting policy, not a promise:

SituationStart hereWhy
Ordinary bounded work with clear checksAny available GPT-5.6 model at medium

Turn AI video into a repeatable engine

Build an AI-assisted video pipeline with hook-first scripts, brand-safe edits, and multi-platform delivery.