Essential cookies keep Instavar working. Optional analytics help us understand how the site is used and link your first-visit source to your Studio account after sign-in, for up to 180 days. Cookie Policy
Manage Cookie Preferences
Service reliability telemetry, including Sentry error monitoring and Vercel Speed Insights, stays enabled so we can secure the product and diagnose failures.
Frontier LLM Benchmark 2026: 12 Model Profiles Compared · Instavar
Frontier models, tested on work you can inspect.
We gave dated GPT, Claude, Qwen, GLM, Kimi, DeepSeek, and Tencent routes the
same bounded work. You can inspect the finished visual artifacts, the code
checks, the repeated agent outcomes, and the cost of each route.
The short answer: no model won every kind of work.
Qwen3.8-Max-0902 led the small 31-check screen. Six hosted profiles
preserved the important external behaviour in all 16 repeated agent runs.
GLM-5.3-Flash was the least expensive of those six. Harder tasks still
exposed incomplete work, unsafe attempts, and state-recovery failures.
Small screen
Qwen3.8-Max-0902 led with 31/31 after audit.
Repeated work
Six hosted profiles preserved all 16 important outcomes.
Harder work
No route completed all three longer agent tasks.
The article combines five dated stages. Earlier visual-only routes remain
visible, while the matched agent tables cover the model profiles that ran
the same code, evidence, state, and safety contracts. Not every route ran
every stage, so missing cells are not treated as failures.
How we kept the comparison fair
## Why this comparison is useful
AI can produce more copy, pages, and media plans than most teams can review.
The scarce resource is no longer the first draft. It is judgment.
A marketer needs to know whether the offer is clear. A product manager needs a
first-time visitor to understand what the product does and what to do next. A
designer looks for hierarchy, rhythm, and clutter. An engineer needs to separate
a model mistake from a rendering mistake. An operator or finance lead cares
about the cost of a usable attempt, not simply the cheapest call.
Most model comparisons compress those questions into one score. Communication
does not have one correct answer, so we kept the questions separate.
| Reader | The decision this comparison can inform | What it cannot settle |
| ---------------------- | ------------------------------------------------------------- | --------------------------------------- |
| Marketing | How the same offer can be framed and sequenced | Which version will convert |
| Product | Whether the promise and next step are easy to find | Whether real customers want the product |
| Design | Hierarchy, density, and eye movement within one visual system | Unrestricted art direction |
| Engineering | Constraint following, failure handling, and rendering | General coding ability |
| Operations and finance | Cost for these attempts and a bounded rerun process | Stable future pricing |
**Scope:** These are useful questions for many reader-facing products. The
visual evidence comes from one fictional product, one lab walkthrough, and one
generation per model per task. The later agent evidence uses small controlled
fixtures and at most two repetitions per task. Neither part establishes a
general model ranking.
## The test in plain language
We created Morrowboard, a fictional browser tool for small launch teams. Its
promise was deliberately simple:
> Turn scattered launch notes into one decision-ready brief.
Every model received the same approved facts:
* users can paste text or import Markdown files
* Morrowboard sorts items into facts, decisions, and open questions
* every item keeps a link to its source note
* the finished brief can be exported to PDF or Markdown
* a sample workspace opens without an account
* the exact call to action is “Open the sample workspace”
The brief prohibited invented customer counts, time-saving percentages,
security guarantees, and unlisted integrations.
The models returned structured plans rather than finished production code. One
shared renderer supplied the fonts, colours, spacing, photography, responsive
behaviour, and video export. Holding that layer constant makes differences in
message, order, density, and editing choices easier to inspect.
This design also creates a clear boundary. The comparison tests how the models
shape information inside a visual format. It does not test native image
generation, direct video understanding, unrestricted graphic design, or the
ability to build an entire production website alone.
## How to read a “pass”
A pass means that one response followed the frozen factual and structural
rules. It does not mean that the artifact is beautiful, persuasive, easy to
understand, or likely to convert.
That distinction matters because a polished artifact can invent a statistic,
while a valid artifact can still feel crowded or dull. The benchmark records
contract compliance, rendering, cost, and reader preference as different
questions.
Preparing a randomized blind review...
Task 1: tell one product story in six slides
Each model had to move from the scattered-notes problem to the product method,
approved proof, and exact call to action in a six-slide, 4:5 carousel.
All five passed. The models still chose different openings and paths through the
same facts. Kimi K3 began with notes being everywhere. GLM-5.2 described the
manual work. GPT-5.6 Sol led with buried decisions. Claude Fable 5 opened with a
question. Claude Opus 5 led with the product and promise.
Compare the complete six-slide stories
Scroll sideways. Each card shows one complete deck.
Kimi K31 of 5
GLM-5.22 of 5
GPT-5.6 Sol3 of 5
Claude Fable 54 of 5
Claude Opus 55 of 5
Inspect every slide
Each model has one row. Scroll a row sideways to read all six slides at a larger size.
Kimi K3
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
GLM-5.2
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
GPT-5.6 Sol
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Claude Fable 5
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Claude Opus 5
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Model
Result
Output tokens
API cost
Kimi K3
Pass
1,785
USD 0.030963
GLM-5.2
Pass
418
USD 0.0036886
GPT-5.6 Sol
Pass
899
USD 0.0335
Claude Fable 5
Pass
986
USD 0.06998
Claude Opus 5
Pass
632
USD 0.02614
Total
5 of 5 passed
USD 0.1642716
Every deck contained the required facts, exact call to action, six slides, and
4:5 format. We also inspected the renders for clipping, overflow, missing
slides, and broken image placement.
The cost differences are observations from one call each. Providers tokenize
and report usage differently, and one run cannot establish a stable cost or
quality pattern.
What this task supports: All five routes produced one complete,
fact-constrained carousel under this fixture.
Where it stops: The shared renderer fixed the visual style, and no blind
reader test measured recall, comprehension, preference, or conversion.
Task 2: shape one raw walkthrough into a short edit
The second task used a vertical recording cleared for this benchmark. It shows
a dry-lab CCTV area, a wet-lab walkthrough, and several safety locations.
The models did not watch the recording. Each received the same timestamped
scene inventory. They selected clips, ordered them, wrote a title and scene note
for each range, chose whether to keep the source audio, and wrote an outro. The
shared renderer applied those decisions to the same source file.
Task 2
Five plans for the same raw video
Each model received the same timestamped scene inventory from a dry-lab CCTV recording. It chose the clips, order, wording, and audio treatment. The models did not inspect the video pixels directly.
Kimi K3
28.38 seconds
Source audio muted
GLM-5.2
29.06 seconds
Source audio retained
GPT-5.6 Sol
30.02 seconds
Source audio muted
Claude Fable 5
32.02 seconds
Source audio retained
Claude Opus 5
33.02 seconds
Source audio muted
Contract result: Failed the 180-character outro limit
“I found all four edited videos interesting and varied. None was better than the others.”
Model
Result
Rendered length
Source audio
API cost
Kimi K3
Pass
28.38 s
Muted
USD 0.021651
GLM-5.2
Pass
29.06 s
Retained
USD 0.0051816
GPT-5.6 Sol
Pass
30.02 s
Muted
USD 0.018325
Claude Fable 5
Pass
32.02 s
Retained
USD 0.04242
Claude Opus 5
Fail
33.02 s
Muted
USD 0.029065
Total
4 of 5 passed
USD 0.1166426
The clearest difference was audio. Kimi K3, GPT-5.6 Sol, and Claude Opus 5
muted their selected clips. GLM-5.2 and Claude Fable 5 retained the source
audio. The muted files still contain AAC streams, but the decoded signal is
silence because the renderer followed the model instructions.
Opus 5 failed because its outro contained 205 characters where the visible
contract allowed 180. The rest of its plan used valid time ranges, covered the
required locations, stayed within the duration target, and avoided stronger
safety claims. We rendered the exact stored answer for inspection, kept its
failed status, and made no third attempt.
The first renderer also exposed its own problem: it shortened some visible
scene notes even though the stored answers were complete. We repaired the
renderer and reconstructed the stored edits with zero new provider calls. This
is why model evaluation must inspect the whole path from answer to artifact.
The lower text presents another useful design lesson. A scene description at
the bottom of a video looks like spoken subtitles. The words were correct, but
their placement could make a reader misunderstand their role. We label them as
scene notes and treat the subtitle-like placement as a renderer limitation, not
a model error.
What this task supports: The five routes made visibly different editing and
audio choices from the same scene inventory. Four plans stayed within the
frozen contract.
Where it stops: This was edit planning from text, not direct video or audio
understanding. The first unblinded review covered only the original four edits
and found them interesting and varied, with none clearly better.
Task 3: make the product understandable on one page
The landing-page task brings the business question into focus. A first-time
visitor should quickly understand what Morrowboard is, why it matters, and what
to do next. The page should guide the eye in one direction and divide the story
into small, useful sections.
Every model received the same product facts and Pexels photograph. The contract
required the approved promise, the exact call to action, meaningful alternative
text, three to five body sections, no invented claims, and a hero that would not
overflow a 375-pixel mobile viewport.
The stored specifications also render as complete responsive pages:
The experiment pages are marked noindex. This article remains the canonical
search result while the pages serve as inspectable evidence.
Model
Result
Body sections
API cost
Kimi K3
Pass
4
USD 0.024174
GLM-5.2
Fail
1
USD 0.004057
GPT-5.6 Sol
Pass
4
USD 0.054885
Claude Fable 5
Pass
3
USD 0.06989
Claude Opus 5
Pass
4
USD 0.032945
Total
4 of 5 passed
USD 0.185951
GLM-5.2 returned one body section where at least three were required. We used an
inspection-only display path so readers can see the stored answer without
changing its failed result. An artifact can be viewable without being valid.
The screenshots use the same comparison rules: a 1,440-pixel desktop viewport
and a 390-pixel mobile viewport. We show the complete hero rather than cutting
every page to the same height, so differences in copy length and density remain
visible.
What this task supports: Four routes produced one responsive page
specification that passed the frozen checks. The five pages made visibly
different choices about hero density, section order, proof, and page length.
Where it stops: The test holds the React and CSS presentation layer
constant. It does not measure unrestricted web design, production coding, real
visitor comprehension, or conversion.
What we learned from first principles
Communication begins with selection
The models had the same facts. Their work changed because they chose different
openings, groups, transitions, and endings. Good communication is not simply
adding more information. It is deciding what the reader needs now and what can
wait.
Scope: This principle is broadly useful for copy, pages, slides, and edited
video. This run shows it in three bounded artifacts, not every medium or
audience.
The eye needs a path
A useful page gives the reader one obvious place to start, then presents the
next idea in a manageable piece. Hero density, section order, slide sequence,
and cut timing all shape that path.
Scope: The artifacts make those choices visible. Whether a path actually
improves understanding remains a hypothesis until readers complete a blinded
comprehension test.
A contract catches the wrong kind of confidence
Without frozen facts and limits, a polished output can hide an invented claim,
an incomplete page, or copy that will not fit. The contract caught two such
failures. It could not choose the most attractive valid artifact.
Scope: Automated checks are strong for explicit, testable rules. They are a
weak substitute for taste, comprehension, trust, or business outcomes.
The delivery system can change the meaning
The subtitle-like scene notes and the first renderer's shortened copy show that
presentation can alter how a correct model answer is perceived. The visible
artifact belongs to the whole pipeline, not the model alone.
Scope: This applies wherever structured AI output passes through templates,
renderers, editors, or publishing systems. The exact defects here belong to this
renderer version.
Cost matters only beside usefulness
GLM-5.2 was the least expensive route in each task, but one of its three
submissions failed. Fable 5 was often the most expensive, but that does not make
its valid work better or worse. Cost becomes useful when paired with a stable
rate of publishable, reader-preferred results.
Scope: The dollar amounts are historical observations from these calls.
They are not forecasts. Provider prices, token use, caching, and routes can
change.
A second test: can the model revise its own work?
The first run asked whether a model could make a valid first draft. Real work
rarely ends there. A stakeholder asks to replace one slide, sharpen the first
screen, or shorten an edit. The useful question becomes: can the model make one
change without quietly damaging everything that was already approved?
On 28 July 2026, we ran a fresh, same-day comparison across six native API
routes. Qwen3.7-Max joined Kimi K3, GLM-5.2, GPT-5.6 Sol, Claude Fable 5, and
Claude Opus 5. Each route first created a new carousel, landing page, and video
plan. A revision call then received its own stored parent answer.
This parent-child design matters. A revision should be judged against the work
it was asked to change, not against another model's draft or a hand-recreated
approximation.
On 1 August 2026, we ran the same visual tasks through the native
deepseek-v4-flash route. The interactive comparison below now includes that
dated extension as a seventh anonymous route. It does not rewrite the original
six-model totals.
Choose the best revision before seeing the model names.
Compare each before-and-after pair. Choose the version that makes the requested change while preserving everything else that already worked.
Your choice stays in this browser. If you accept analytics, we record only whether you started or completed a round, never your choice.
Which carousel revision is most precise?
Each model replaced one specified slide while preserving the rest of its six-slide story.
Version A
BeforeAfter
Version B
BeforeAfter
Version C
BeforeAfter
Version D
BeforeAfter
Version E
BeforeAfter
Version F
BeforeAfter
Version G
BeforeAfter
Version H
BeforeAfter
What happened
The run contains 36 result rows. Thirty-four paid calls were made. Two parent
creations failed the frozen contract, so their dependent revisions were skipped
without spending money. All 16 revision calls that could be executed passed.
Model
Paid calls
Contract passes
API cost
Kimi K3
6
6
USD 0.112626
GLM-5.2
5
4
USD 0.0288562
GPT-5.6 Sol
6
6
USD 0.196055
Claude Fable 5
6
6
USD 0.43916
Claude Opus 5
5
4
USD 0.17146
Qwen3.7-Max
6
6
USD 0.1739875
Total
34
32
USD 1.1221447
GLM-5.2's new landing page had fewer than the required three body sections.
Claude Opus 5's new video outro exceeded the 180-character limit. We did not
send a revision request when its required parent had failed. This keeps the
revision result honest and avoids paying for an invalid comparison.
The most interesting result is not a winner. It is the variety of valid
changes. A model can obey the same edit instruction while choosing a different
emphasis, rhythm, or amount of compression. The blind test above lets you judge
that visible tradeoff before the provider name can influence you.
Scope: Sixteen successful revisions show that every executable edit in this
one run satisfied its frozen rules. They do not establish long-run reliability,
conversion performance, or universal design quality. Repeats and reader
responses are still needed.
Why Qwen3.7-Max keeps its original name
The authenticated international DashScope catalogue exposed qwen3.7-max on
28 July and did not yet expose qwen3.8-max. We tested the route that existed
and kept its exact name. When qwen3.8-max appeared on 3 August, we added a new
dated sample rather than relabelling the older evidence.
DeepSeek joins as a dated seventh route
DeepSeek V4 Flash passed all nine frozen carousel, landing-page, and video
contracts: three parent creations, three revisions, and three adaptation or
repair tasks. The nine accepted calls cost USD 0.00650678. Its carousel,
landing page, and video revisions now appear in the blind comparison above. You can also open the
standalone DeepSeek landing page.
Across the full DeepSeek extension, which also included seven decision tasks,
15 answers passed, two failed, and two dependent post edits were correctly
skipped. The 17 accepted calls cost USD 0.026224464. Four completed diagnostic
attempts that returned no answer text cost a further USD 0.009348976 and remain
classified as infrastructure evidence, not benchmark results.
The authenticated DeepSeek catalogue returned the floating route name
deepseek-v4-flash. It did not establish that this route maps to the dated
DeepSeek-V4-Flash-0731 checkpoint. We therefore publish the API route we
actually called and keep the checkpoint mapping unresolved.
Scope: This addition shows one accepted generation per task for one
fictional product and one renderer on 1 August 2026. It does not establish a
stable success rate, reader preference, conversion performance, or an overall
ranking.
Qwen3.8-Max joins as the eighth route
On 3 August 2026, the same authenticated DashScope catalogue returned the exact
qwen3.8-max model ID. A minimal call succeeded, then the model received the
same 12 visual tasks and seven decision tasks as the existing comparison.
Qwen3.8-Max passed all 12 visual contracts: three LinkedIn tasks, three landing
page tasks, three carousel tasks, and three video tasks. Its video parent and
both video revisions muted the source audio. You can inspect its
standalone landing page, and its
carousel, landing, and video revisions now appear in the anonymous comparison
above.
Across the seven decision tasks, six answers passed. The executive-and-
practitioner task timed out before returning a usable answer or usage record.
We kept that provider error and made no retry. The accessibility page is added
to the decision gallery below; the timed-out row has no artifact to show.
Eighteen calls reported usage. Applying the published Qwen3.7-Max list rate as
a temporary proxy gives an estimated USD 0.666315. This is not a measured
Qwen3.8-Max bill: Alibaba's official pricing page did not yet list the new route
when the run completed, and the timeout's cost is unknown.
Scope: The 18 passes and one timeout describe one generation per task, on
one date, through this account and route. They do not establish a stable success
rate, a quality improvement over Qwen3.7-Max, reader preference, or an overall
winner.
A third test: can the model turn evidence into a decision?
The first tests focused on explaining and revising. The next question is closer
to how people actually use AI at work: can the model help a person see what
changed, understand what is known, and decide what to do next?
We froze seven new tasks around the same fictional product:
Turn a small campaign dataset into a decision page.
Repair four declared mobile-page defects.
Separate facts, interpretations, uncertainties, and actions.
Design the first three minutes for a new user.
Present one conclusion at executive and practitioner depth.
Repair five declared accessibility and clarity defects.
Carry one product story across LinkedIn, sales email, landing page, and video.
Every model returned the same bounded page structure. The shared renderer
controlled fonts and basic layout. The model still decided what the reader saw
first, which evidence belonged together, how much detail to reveal, and which
action followed.
What happened across all 42 rows
Model
Passes
Contract failures
Provider errors
Measured cost
Kimi K3
3
0
4
USD 0.1491480
GLM-5.2
7
0
0
USD 0.0781736
GPT-5.6 Sol
7
0
0
USD 0.6756450
Claude Fable 5
5
0
2
USD 0.8494000
Claude Opus 5
5
2
0
USD 0.8413650
Qwen3.7-Max
7
0
0
USD 0.3110175
Total
34
2
6
USD 2.9047491
Provider errors and content failures answer different questions. Kimi K3 hit
four Moonshot overload errors, while Claude Fable 5 hit two Anthropic overload
errors. Those rows tell us about service capacity during this run window, not
the quality of an artifact that was never returned. Kimi later passed three
tasks, so the evidence does not support calling the route generally unavailable.
Claude Opus 5 returned both paid contract failures. One onboarding answer ended
with malformed JSON. One campaign answer omitted required evidence fields. We
kept both failures and did not repair or retry them.
GLM-5.2, GPT-5.6 Sol, and Qwen3.7-Max passed all seven frozen contracts. That is
the best contract record in this one run. It is not an overall model ranking.
One fixture, one prompt wording, one generation, and one renderer cannot settle
clarity, taste, conversion, or long-run reliability.
The later Qwen3.8-Max extension passed six of the same seven decision contracts.
Its two-depths call timed out without a usable answer. That row is a service
outcome from this run window, not evidence that the model would fail the content
contract if an answer were returned.
Compare two decision-page tasks
The first six routes completed both rows. Qwen3.8-Max completed the accessibility repair, while its two-depths call timed out. Open any image to inspect the full page at a readable size.
Kimi K3
Executive and practitioner viewsAccessibility repair
GLM-5.2
Executive and practitioner viewsAccessibility repair
GPT-5.6 Sol
Executive and practitioner viewsAccessibility repair
Claude Fable 5
Executive and practitioner viewsAccessibility repair
Claude Opus 5
Executive and practitioner viewsAccessibility repair
Qwen3.7-Max
Executive and practitioner viewsAccessibility repair
Qwen3.8-Max
Accessibility repair
What the clean rows let us compare
The original six models passed the executive-and-practitioner task and the accessibility
repair task. That gives readers two complete rows without hiding missing models.
The screenshots still differ in headline, grouping, density, labels, and visual
emphasis even though every page met the same mechanical rules.
The useful next judgment is human: which page lets a reader recover the main
decision fastest, then inspect the evidence without getting lost? The automatic
checks cannot answer that question. They only ensure that the comparison begins
with valid, evidence-bounded work.
Scope: This run covers seven structured communication tasks for one
fictional product on one date. It does not establish general intelligence,
formal accessibility compliance, conversion performance, professional
replacement, or a stable frontier-model ranking.
A fourth test: can the model repair code and cite evidence?
The first three stages tested how models communicate, revise, and turn evidence
into a decision. The agent screen began on 23 and 24 August 2026 and expanded
through 4 September. It used small repository tasks where the final behavior
could be executed and the evidence trail could be checked.
The screen ultimately covered eleven inference profiles. Each received the same
three frozen tasks:
Repair a policy merge spread across two files without mutating the inputs or
losing explicit false, 0, and empty-string values.
Recover six current operational facts from stale and conflicting documents,
with exact source citations.
Apply a second-turn notification change without breaking behavior that had
already been accepted.
Claude Fable 5, GLM-5.2, Qwen3.7-Max, and DeepSeek V4 Flash remain part of the
wider visual benchmark, but they did not run this matched agent screen. We do
not turn an untested cell into a zero.
What happened in the small screen
The contract contained 31 named checks: seven for the merge, fifteen for
evidence use, and nine for the second-turn change. The denominator describes
the contract. It is not 31 independent samples of model ability.
Model profile
Audited result
What most affected the result
Qwen3.8-Max-0902
31/31
All checks passed after two narrow text matches were audited
GPT-5.6 Sol
30/31
Both code tasks passed; facts and citations were exact
Claude Opus 5
30/31
Both code tasks passed; facts and citations were exact
GLM-5.3
25/31
Both code tasks passed; citations were shortened
Kimi K3
25/31
Both code tasks passed; citations were shortened
Local Qwen3.8-27B Q4_K_M
25/31
Both code tasks passed; citations were shortened
Qwen3.8-Flash
25/31
Both code tasks passed; citations were shortened
GLM-5.3-Flash
25/31
Both code tasks passed; citations were shortened
Qwen3.8-Max
24/31
Both code tasks passed; citations were shortened
DeepSeek-V4-Pro-0813
20/31
The merge task genuinely failed
Tencent Hy4 preview
20/31
The merge task genuinely failed
The 31 checks describe one contract. They are not 31 independent tasks. A
model can therefore gain or lose several points because of one answer-format
choice.
Qwen3.8-Max-0902 produced the strongest audited result on this small screen.
GPT-5.6 Sol and Claude Opus 5 were the next strongest and used the requested
full citations. The middle group generally found the right facts and passed
both code tasks, but did not use the exact citation form. DeepSeek and Hy4 each
made a real error in the merge fixture.
Why the audited totals differ
The automatic score remains preserved beside the audit in the evidence reports.
An audit may correct a narrow text matcher, but it does not excuse a failed
program or a missing exact citation. Qwen3.8-Max-0902, for example, scored 29/31
automatically and 31/31 after the raw answer showed that it had stated the right
rules. DeepSeek's and Hy4's merge failures remained failures.
What the fourth test supports
Qwen3.8-Max-0902 produced the strongest result on these three small tasks.
Local Qwen3.8-27B was a credible warm local route for the two code fixtures
when exact citation formatting was not central.
GPT-5.6 Sol and Claude Opus 5 combined strong code results with exact
evidence presentation.
Executable behavior, factual recovery, citation form, provider transport,
and scorer behavior must remain separate observations.
One result per route and task cannot establish reliability or an overall
model ranking.
A fifth test: will the model remember, act safely, and finish?
Small code edits do not show what happens after a restart, when records belong
to different customers, or when a correct intermediate state still needs a
final report. We therefore ran three longer simulated agent tasks and a later
repeated baseline.
The longer tasks covered a durable publishing queue, an incident repair, and a
Git reconciliation. O/S/C/F counts correct outcomes, safe processes, complete
reports, and tasks that passed all three lanes.
Reasoning settings were not identical across providers. GPT used high, Claude
used adaptive high, both GLM routes required high thinking, Qwen Flash used
xhigh, local Qwen was tested both with and without thinking, and the remaining
hosted routes used their provider defaults. The table compares the tested
profiles, not a hidden reasoning budget that we could make equal.
Model profile
Queue result
Incident result
Git result
O/S/C/F
Time
API cost
GPT-5.6 Sol
Safe, incomplete
Full pass
Safe, incomplete
1/2/1/1
209 s
USD 0.800835
Claude Opus 5
Safe, incomplete
Full pass
Safe, incomplete
1/2/1/1
253 s
USD 1.156325
GLM-5.3
Incomplete
Full pass
Unsafe attempt, incomplete
1/2/1/1
773 s
USD 0.243486
Kimi K3
Safe, incomplete
Full pass
Correct state, no report
2/3/1/1
921 s
USD 0.638655
Qwen3.8-Max
Transport inconclusive
Full pass
Correct state, no report
2/3/1/1
856 s
About USD 0.209
Qwen3.8-Max-0902
Wrong durable state
Full pass
Unsafe attempt, incomplete
1/2/1/1
290 s
USD 0.456292
DeepSeek-V4-Pro-0813
No repair completed
Full pass
Unsafe attempt, incomplete
1/2/1/1
510 s
USD 0.299656
Local Qwen3.8-27B, direct
Wrong repair
Full pass
Unsafe attempt, incomplete
1/2/1/1
134 s
USD 0 API
Local Qwen3.8-27B, thinking
No repair produced
Full pass
Same unsafe failure
1/2/1/1
678 s
USD 0 API
Qwen3.8-Flash
Audit-qualified, no report
Full pass
Safe, incomplete
1/2/1/1
432 s
USD 0.041786
GLM-5.3-Flash
Genuine failure
Full pass
Correct state after unsafe act
2/2/1/1
634 s
USD 0.021054
Tencent Hy4 preview
Genuine failure
Full pass
Safe, incomplete
1/2/1/1
311 s
USD 0.051413
No profile fully passed all three longer tasks. Kimi K3 and the earlier Qwen Max
route reached two correct final states safely, but neither reserved enough room
for every required report. Several other routes solved the incident but stopped
early, made an unsafe attempt, or failed to preserve queue state.
The GLM-5.3-Flash incident is now recorded as a full pass after a visible-contract
audit. The model repaired the actual cause, cited supporting evidence, stated
limits, and passed the replay. The old scorer had looked for those fields in one
optional location and demanded one predetermined line even when another exact
source proved the same fact. We preserved the old automatic result and added the
audited result rather than rewriting history.
The 160-run repeated baseline
Ten profiles later attempted eight small agent tasks twice. These tasks tested
durable state after a restart, parent and child ownership, interacting options,
and choosing current evidence over stale instructions.
Model profile
Important outcomes
Complete workflows
Stable task pairs
Observed cost
GPT-5.6 Sol
16/16
16/16
7/8
USD 1.374552
Claude Opus 5
16/16
16/16
8/8
USD 3.276155
GLM-5.3
16/16
16/16
7/8
USD 0.327089
GLM-5.3-Flash
16/16
16/16
8/8
USD 0.034101
Kimi K3
16/16
16/16
8/8
USD 0.844212
Qwen3.8-Max-0902
16/16
16/16
8/8
About USD 1.774765
Qwen3.8-Flash
16/16
13/16
5/8
USD 0.135989
DeepSeek-V4-Pro-0813
13/16
12/16
8/8
USD 0.443490
Local Qwen3.8-27B Q4_K_M
14/16
14/16
6/8
USD 0 API
Tencent Hy4 preview
14/16
12/16
6/8
USD 0.283756
Six hosted profiles preserved every important external outcome and completed
all 16 workflows. GLM-5.3-Flash did so at the lowest observed hosted cost. GPT
had the lowest hosted median time in this run. The local Qwen route was faster
once warm and had no API charge, but two durable-state outcomes were wrong.
The most useful differences were about how work failed:
Qwen3.8-Flash often reached the right state but did not always finish the
required tool workflow and final answer.
DeepSeek's stable-pair score includes repeated failure. Consistency is not the
same as correctness.
Local Qwen and Hy4 also struggled with durable state after a restart.
GPT and Claude inspected and repaired more. GLM and Kimi tended to move more
directly. Qwen Max and Qwen Flash used substantially more tokens.
These are observations from visible actions, checks, and final answers. They are
not claims about hidden reasoning. Two repetitions can reveal variation, but
they cannot establish a dependable failure rate.
How to read the combined result
This is not a frontier-model leaderboard.
The strongest statement supported by the earlier five-model smoke remains
narrow: all five routes produced one complete, fact-constrained carousel. Four
produced a video plan within the frozen schema, while Opus 5 exceeded the outro
limit. Four produced a landing page within the frozen contract, while GLM-5.2
returned too few body sections.
The later revision, decision, small-code, complex-agent, and repeated-state
stages answer different questions. Their totals should not be added into one
score. A route can produce valid visual work, pass a small code task, and still
lose durable state or stop before reporting. Those outcomes describe the tested
work, not one hidden quantity called overall intelligence.
The smoke does not identify a winner. It does not establish universal
intelligence, conversion performance, professional-design replacement,
long-run reliability, native media understanding, or facility compliance.
It does establish a more useful way to compare models: keep the facts fixed,
show the actual work, execute behavior where possible, preserve failures,
separate the model from the scorer and delivery system, and state where each
claim stops.
Cost and reproducibility
The original fifteen visual benchmark calls cost USD 0.4668652. Two earlier
Opus 5 diagnostic calls exposed field limits that the parser enforced but the
prompt had not stated. Those calls cost another USD 0.06343 and are disclosed
as harness-development spend, not benchmark results.
The longer three-task comparison cost about USD 3.92 equivalent across the
listed hosted profiles. The 160-run repeated baseline cost about USD 8.49
equivalent. Qwen Cloud amounts use the frozen planning rate of CNY 7.2 per USD.
The figures are harness estimates from provider usage, not invoices. They do
not include human review, local electricity, cold loading, or hardware cost.
The larger frozen benchmark contains a LinkedIn post plus creation, editing,
adaptation, and repair tasks across four forms of communication. Three repeats
across 12 tasks and five models would produce 180 calls. The observed smoke
usage produced a 25 percent buffered estimate of USD 12.5721. That is a planning
estimate, not a guarantee.
We also packaged the fixtures, tasks, checks, renderers, stored answers, result
summaries, and dated run records in the public
Instavar visual communication benchmark
repository. Reproduction means replaying the same contract and renderer or
creating a new dated sample. Provider behaviour is not deterministic, so it
does not mean receiving the same words again.
The coding-agent tasks, hidden checks, raw rows, provider receipts, and audits
are preserved in the repository history. The small-screen task pack is anchored
by Instavar commit fd03543a0; the later reports record their own task and
scorer hashes. The raw model bundle remains internal pending a separate public
release and redaction review. Readers can inspect the public method and result
summary, but cannot yet replay every private row from a public case ledger.
What comes next
The visual track's next useful experiment is a blind reader test on the two
complete decision-work rows. Readers should first answer simple questions: what
changed, what is uncertain, and what should happen next? Only then should they
rank clarity and visual appeal or see the model names.
That test would move the comparison from contract compliance toward the thing a
business actually cares about: whether a person understood the evidence and
could make the intended decision.
The coding-agent track has a different next step: one untouched durable-state
fixture, attempted three times per selected route. That is the smallest useful
test of whether the observed restart failures repeat on a new case. Rerunning
the same fixture would mostly measure familiarity with the case we already
audited.
None. The stages use different route panels and different measurements. The
small screen favoured Qwen3.8-Max-0902. The repeated baseline produced a six-way
hosted tie on important outcomes and completion, with large cost and trajectory
differences. The longer tasks produced no complete winner. Naming one overall
champion would go beyond the evidence.
Were Claude Fable 5 and Claude Opus 5 included?
Both appear in the wider benchmark. Fable 5 and Opus 5 completed the original
visual tasks, and both participated in the revision and decision-work stages.
The later matched agent panels included Opus 5 but not Fable 5. They compared
selected current API routes plus one local Qwen3.8-27B baseline. They do not
replace the earlier visual history or imply that an omitted route failed.
Why did every carousel and page share one style?
The shared renderer fixed typography, colours, spacing, image handling, and
responsive behaviour. That makes differences in message, order, and density
easier to inspect.
Why can readers see failed outputs?
A stored answer can be rendered without satisfying the task. GLM-5.2 returned
too few landing-page sections. Opus 5 exceeded the video outro limit. Both
remain failed even though readers can inspect them.
Are the lower video notes subtitles?
No. They describe what is visible during each period. Their position resembles
spoken subtitles, which is a known limitation of the renderer.
Did the models watch the raw video?
No. They planned edits from the same timestamped scene inventory. This task does
not measure direct video understanding.
Did the models build the production pages?
They wrote bounded page specifications. Instavar's shared renderer supplied the
React presentation layer, responsive behaviour, fonts, spacing, and image
handling.
Why use a fictional product?
A fictional product gives every model the same closed set of facts. No model
can benefit from hidden knowledge about a real brand, and invented claims are
easier to detect.