Essential cookies keep Instavar working. Optional analytics help us understand how the site is used. Cookie Policy
Manage Cookie Preferences
Service reliability telemetry, including Sentry error monitoring and Vercel Speed Insights, stays enabled so we can secure the product and diagnose failures.
Kimi K3 vs GLM-5.2 vs GPT-5.6 Sol: Visual Benchmark · Instavar
Eight frontier models. One product story.
Kimi K3, GLM-5.2, GPT-5.6 Sol, and five other model routes each explained
the same fictional product as a carousel, video, and landing page. The
finished work makes their different choices easy to see.
The short answer: no model won overall. Most outputs
followed the brief, but they guided the eye, ordered the story, and handled
sound in noticeably different ways. GLM also returned too little for one
landing-page task. This is one run per task, not a universal ranking.
Carousel
Both passed. Different openings.
Video
Both passed. Kimi muted the clip; GLM kept its sound.
Landing page
Kimi passed. GLM returned too little.
Eight routes tested: Kimi K3, GLM-5.2, GPT-5.6 Sol, Claude Fable 5, Claude
Opus 5, Qwen3.7-Max, DeepSeek V4 Flash, and Qwen3.8-Max. Not every route
appears in every run.
How we kept the comparison fair
## Why this comparison is useful
AI can produce more copy, pages, and media plans than most teams can review.
The scarce resource is no longer the first draft. It is judgment.
A marketer needs to know whether the offer is clear. A product manager needs a
first-time visitor to understand what the product does and what to do next. A
designer looks for hierarchy, rhythm, and clutter. An engineer needs to separate
a model mistake from a rendering mistake. An operator or finance lead cares
about the cost of a usable attempt, not simply the cheapest call.
Most model comparisons compress those questions into one score. Communication
does not have one correct answer, so we kept the questions separate.
| Reader | The decision this comparison can inform | What it cannot settle |
| ---------------------- | ------------------------------------------------------------- | --------------------------------------- |
| Marketing | How the same offer can be framed and sequenced | Which version will convert |
| Product | Whether the promise and next step are easy to find | Whether real customers want the product |
| Design | Hierarchy, density, and eye movement within one visual system | Unrestricted art direction |
| Engineering | Constraint following, failure handling, and rendering | General coding ability |
| Operations and finance | Cost for these attempts and a bounded rerun process | Stable future pricing |
**Scope:** These are useful questions for many reader-facing products. The
evidence in this article comes from one fictional product, one lab walkthrough,
one generation per model per task, and eight dated API routes.
## The test in plain language
We created Morrowboard, a fictional browser tool for small launch teams. Its
promise was deliberately simple:
> Turn scattered launch notes into one decision-ready brief.
Every model received the same approved facts:
* users can paste text or import Markdown files
* Morrowboard sorts items into facts, decisions, and open questions
* every item keeps a link to its source note
* the finished brief can be exported to PDF or Markdown
* a sample workspace opens without an account
* the exact call to action is “Open the sample workspace”
The brief prohibited invented customer counts, time-saving percentages,
security guarantees, and unlisted integrations.
The models returned structured plans rather than finished production code. One
shared renderer supplied the fonts, colours, spacing, photography, responsive
behaviour, and video export. Holding that layer constant makes differences in
message, order, density, and editing choices easier to inspect.
This design also creates a clear boundary. The comparison tests how the models
shape information inside a visual format. It does not test native image
generation, direct video understanding, unrestricted graphic design, or the
ability to build an entire production website alone.
## How to read a “pass”
A pass means that one response followed the frozen factual and structural
rules. It does not mean that the artifact is beautiful, persuasive, easy to
understand, or likely to convert.
That distinction matters because a polished artifact can invent a statistic,
while a valid artifact can still feel crowded or dull. The benchmark records
contract compliance, rendering, cost, and reader preference as different
questions.
Preparing a randomized blind review...
Task 1: tell one product story in six slides
Each model had to move from the scattered-notes problem to the product method,
approved proof, and exact call to action in a six-slide, 4:5 carousel.
All five passed. The models still chose different openings and paths through the
same facts. Kimi K3 began with notes being everywhere. GLM-5.2 described the
manual work. GPT-5.6 Sol led with buried decisions. Claude Fable 5 opened with a
question. Claude Opus 5 led with the product and promise.
Compare the complete six-slide stories
Scroll sideways. Each card shows one complete deck.
Kimi K31 of 5
GLM-5.22 of 5
GPT-5.6 Sol3 of 5
Claude Fable 54 of 5
Claude Opus 55 of 5
Inspect every slide
Each model has one row. Scroll a row sideways to read all six slides at a larger size.
Kimi K3
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
GLM-5.2
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
GPT-5.6 Sol
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Claude Fable 5
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Claude Opus 5
Slide 1
Slide 2
Slide 3
Slide 4
Slide 5
Slide 6
Model
Result
Output tokens
API cost
Kimi K3
Pass
1,785
USD 0.030963
GLM-5.2
Pass
418
USD 0.0036886
GPT-5.6 Sol
Pass
899
USD 0.0335
Claude Fable 5
Pass
986
USD 0.06998
Claude Opus 5
Pass
632
USD 0.02614
Total
5 of 5 passed
USD 0.1642716
Every deck contained the required facts, exact call to action, six slides, and
4:5 format. We also inspected the renders for clipping, overflow, missing
slides, and broken image placement.
The cost differences are observations from one call each. Providers tokenize
and report usage differently, and one run cannot establish a stable cost or
quality pattern.
What this task supports: All five routes produced one complete,
fact-constrained carousel under this fixture.
Where it stops: The shared renderer fixed the visual style, and no blind
reader test measured recall, comprehension, preference, or conversion.
Task 2: shape one raw walkthrough into a short edit
The second task used a vertical recording cleared for this benchmark. It shows
a dry-lab CCTV area, a wet-lab walkthrough, and several safety locations.
The models did not watch the recording. Each received the same timestamped
scene inventory. They selected clips, ordered them, wrote a title and scene note
for each range, chose whether to keep the source audio, and wrote an outro. The
shared renderer applied those decisions to the same source file.
Task 2
Five plans for the same raw video
Each model received the same timestamped scene inventory from a dry-lab CCTV recording. It chose the clips, order, wording, and audio treatment. The models did not inspect the video pixels directly.
Kimi K3
28.38 seconds
Source audio muted
GLM-5.2
29.06 seconds
Source audio retained
GPT-5.6 Sol
30.02 seconds
Source audio muted
Claude Fable 5
32.02 seconds
Source audio retained
Claude Opus 5
33.02 seconds
Source audio muted
Contract result: Failed the 180-character outro limit
“I found all four edited videos interesting and varied. None was better than the others.”
Model
Result
Rendered length
Source audio
API cost
Kimi K3
Pass
28.38 s
Muted
USD 0.021651
GLM-5.2
Pass
29.06 s
Retained
USD 0.0051816
GPT-5.6 Sol
Pass
30.02 s
Muted
USD 0.018325
Claude Fable 5
Pass
32.02 s
Retained
USD 0.04242
Claude Opus 5
Fail
33.02 s
Muted
USD 0.029065
Total
4 of 5 passed
USD 0.1166426
The clearest difference was audio. Kimi K3, GPT-5.6 Sol, and Claude Opus 5
muted their selected clips. GLM-5.2 and Claude Fable 5 retained the source
audio. The muted files still contain AAC streams, but the decoded signal is
silence because the renderer followed the model instructions.
Opus 5 failed because its outro contained 205 characters where the visible
contract allowed 180. The rest of its plan used valid time ranges, covered the
required locations, stayed within the duration target, and avoided stronger
safety claims. We rendered the exact stored answer for inspection, kept its
failed status, and made no third attempt.
The first renderer also exposed its own problem: it shortened some visible
scene notes even though the stored answers were complete. We repaired the
renderer and reconstructed the stored edits with zero new provider calls. This
is why model evaluation must inspect the whole path from answer to artifact.
The lower text presents another useful design lesson. A scene description at
the bottom of a video looks like spoken subtitles. The words were correct, but
their placement could make a reader misunderstand their role. We label them as
scene notes and treat the subtitle-like placement as a renderer limitation, not
a model error.
What this task supports: The five routes made visibly different editing and
audio choices from the same scene inventory. Four plans stayed within the
frozen contract.
Where it stops: This was edit planning from text, not direct video or audio
understanding. The first unblinded review covered only the original four edits
and found them interesting and varied, with none clearly better.
Task 3: make the product understandable on one page
The landing-page task brings the business question into focus. A first-time
visitor should quickly understand what Morrowboard is, why it matters, and what
to do next. The page should guide the eye in one direction and divide the story
into small, useful sections.
Every model received the same product facts and Pexels photograph. The contract
required the approved promise, the exact call to action, meaningful alternative
text, three to five body sections, no invented claims, and a hero that would not
overflow a 375-pixel mobile viewport.
The stored specifications also render as complete responsive pages:
The experiment pages are marked noindex. This article remains the canonical
search result while the pages serve as inspectable evidence.
Model
Result
Body sections
API cost
Kimi K3
Pass
4
USD 0.024174
GLM-5.2
Fail
1
USD 0.004057
GPT-5.6 Sol
Pass
4
USD 0.054885
Claude Fable 5
Pass
3
USD 0.06989
Claude Opus 5
Pass
4
USD 0.032945
Total
4 of 5 passed
USD 0.185951
GLM-5.2 returned one body section where at least three were required. We used an
inspection-only display path so readers can see the stored answer without
changing its failed result. An artifact can be viewable without being valid.
The screenshots use the same comparison rules: a 1,440-pixel desktop viewport
and a 390-pixel mobile viewport. We show the complete hero rather than cutting
every page to the same height, so differences in copy length and density remain
visible.
What this task supports: Four routes produced one responsive page
specification that passed the frozen checks. The five pages made visibly
different choices about hero density, section order, proof, and page length.
Where it stops: The test holds the React and CSS presentation layer
constant. It does not measure unrestricted web design, production coding, real
visitor comprehension, or conversion.
What we learned from first principles
Communication begins with selection
The models had the same facts. Their work changed because they chose different
openings, groups, transitions, and endings. Good communication is not simply
adding more information. It is deciding what the reader needs now and what can
wait.
Scope: This principle is broadly useful for copy, pages, slides, and edited
video. This run shows it in three bounded artifacts, not every medium or
audience.
The eye needs a path
A useful page gives the reader one obvious place to start, then presents the
next idea in a manageable piece. Hero density, section order, slide sequence,
and cut timing all shape that path.
Scope: The artifacts make those choices visible. Whether a path actually
improves understanding remains a hypothesis until readers complete a blinded
comprehension test.
A contract catches the wrong kind of confidence
Without frozen facts and limits, a polished output can hide an invented claim,
an incomplete page, or copy that will not fit. The contract caught two such
failures. It could not choose the most attractive valid artifact.
Scope: Automated checks are strong for explicit, testable rules. They are a
weak substitute for taste, comprehension, trust, or business outcomes.
The delivery system can change the meaning
The subtitle-like scene notes and the first renderer's shortened copy show that
presentation can alter how a correct model answer is perceived. The visible
artifact belongs to the whole pipeline, not the model alone.
Scope: This applies wherever structured AI output passes through templates,
renderers, editors, or publishing systems. The exact defects here belong to this
renderer version.
Cost matters only beside usefulness
GLM-5.2 was the least expensive route in each task, but one of its three
submissions failed. Fable 5 was often the most expensive, but that does not make
its valid work better or worse. Cost becomes useful when paired with a stable
rate of publishable, reader-preferred results.
Scope: The dollar amounts are historical observations from these calls.
They are not forecasts. Provider prices, token use, caching, and routes can
change.
A second test: can the model revise its own work?
The first run asked whether a model could make a valid first draft. Real work
rarely ends there. A stakeholder asks to replace one slide, sharpen the first
screen, or shorten an edit. The useful question becomes: can the model make one
change without quietly damaging everything that was already approved?
On 28 July 2026, we ran a fresh, same-day comparison across six native API
routes. Qwen3.7-Max joined Kimi K3, GLM-5.2, GPT-5.6 Sol, Claude Fable 5, and
Claude Opus 5. Each route first created a new carousel, landing page, and video
plan. A revision call then received its own stored parent answer.
This parent-child design matters. A revision should be judged against the work
it was asked to change, not against another model's draft or a hand-recreated
approximation.
On 1 August 2026, we ran the same visual tasks through the native
deepseek-v4-flash route. The interactive comparison below now includes that
dated extension as a seventh anonymous route. It does not rewrite the original
six-model totals.
Choose the best revision before seeing the model names.
Compare each before-and-after pair. Choose the version that makes the requested change while preserving everything else that already worked.
Your choice stays in this browser. If you accept analytics, we record only whether you started or completed a round, never your choice.
Which carousel revision is most precise?
Each model replaced one specified slide while preserving the rest of its six-slide story.
Version A
BeforeAfter
Version B
BeforeAfter
Version C
BeforeAfter
Version D
BeforeAfter
Version E
BeforeAfter
Version F
BeforeAfter
Version G
BeforeAfter
Version H
BeforeAfter
What happened
The run contains 36 result rows. Thirty-four paid calls were made. Two parent
creations failed the frozen contract, so their dependent revisions were skipped
without spending money. All 16 revision calls that could be executed passed.
Model
Paid calls
Contract passes
API cost
Kimi K3
6
6
USD 0.112626
GLM-5.2
5
4
USD 0.0288562
GPT-5.6 Sol
6
6
USD 0.196055
Claude Fable 5
6
6
USD 0.43916
Claude Opus 5
5
4
USD 0.17146
Qwen3.7-Max
6
6
USD 0.1739875
Total
34
32
USD 1.1221447
GLM-5.2's new landing page had fewer than the required three body sections.
Claude Opus 5's new video outro exceeded the 180-character limit. We did not
send a revision request when its required parent had failed. This keeps the
revision result honest and avoids paying for an invalid comparison.
The most interesting result is not a winner. It is the variety of valid
changes. A model can obey the same edit instruction while choosing a different
emphasis, rhythm, or amount of compression. The blind test above lets you judge
that visible tradeoff before the provider name can influence you.
Scope: Sixteen successful revisions show that every executable edit in this
one run satisfied its frozen rules. They do not establish long-run reliability,
conversion performance, or universal design quality. Repeats and reader
responses are still needed.
Why Qwen3.7-Max keeps its original name
The authenticated international DashScope catalogue exposed qwen3.7-max on
28 July and did not yet expose qwen3.8-max. We tested the route that existed
and kept its exact name. When qwen3.8-max appeared on 3 August, we added a new
dated sample rather than relabelling the older evidence.
DeepSeek joins as a dated seventh route
DeepSeek V4 Flash passed all nine frozen carousel, landing-page, and video
contracts: three parent creations, three revisions, and three adaptation or
repair tasks. The nine accepted calls cost USD 0.00650678. Its carousel,
landing page, and video revisions now appear in the blind comparison above. You can also open the
standalone DeepSeek landing page.
Across the full DeepSeek extension, which also included seven decision tasks,
15 answers passed, two failed, and two dependent post edits were correctly
skipped. The 17 accepted calls cost USD 0.026224464. Four completed diagnostic
attempts that returned no answer text cost a further USD 0.009348976 and remain
classified as infrastructure evidence, not benchmark results.
The authenticated DeepSeek catalogue returned the floating route name
deepseek-v4-flash. It did not establish that this route maps to the dated
DeepSeek-V4-Flash-0731 checkpoint. We therefore publish the API route we
actually called and keep the checkpoint mapping unresolved.
Scope: This addition shows one accepted generation per task for one
fictional product and one renderer on 1 August 2026. It does not establish a
stable success rate, reader preference, conversion performance, or an overall
ranking.
Qwen3.8-Max joins as the eighth route
On 3 August 2026, the same authenticated DashScope catalogue returned the exact
qwen3.8-max model ID. A minimal call succeeded, then the model received the
same 12 visual tasks and seven decision tasks as the existing comparison.
Qwen3.8-Max passed all 12 visual contracts: three LinkedIn tasks, three landing
page tasks, three carousel tasks, and three video tasks. Its video parent and
both video revisions muted the source audio. You can inspect its
standalone landing page, and its
carousel, landing, and video revisions now appear in the anonymous comparison
above.
Across the seven decision tasks, six answers passed. The executive-and-
practitioner task timed out before returning a usable answer or usage record.
We kept that provider error and made no retry. The accessibility page is added
to the decision gallery below; the timed-out row has no artifact to show.
Eighteen calls reported usage. Applying the published Qwen3.7-Max list rate as
a temporary proxy gives an estimated USD 0.666315. This is not a measured
Qwen3.8-Max bill: Alibaba's official pricing page did not yet list the new route
when the run completed, and the timeout's cost is unknown.
Scope: The 18 passes and one timeout describe one generation per task, on
one date, through this account and route. They do not establish a stable success
rate, a quality improvement over Qwen3.7-Max, reader preference, or an overall
winner.
A third test: can the model turn evidence into a decision?
The first tests focused on explaining and revising. The next question is closer
to how people actually use AI at work: can the model help a person see what
changed, understand what is known, and decide what to do next?
We froze seven new tasks around the same fictional product:
Turn a small campaign dataset into a decision page.
Repair four declared mobile-page defects.
Separate facts, interpretations, uncertainties, and actions.
Design the first three minutes for a new user.
Present one conclusion at executive and practitioner depth.
Repair five declared accessibility and clarity defects.
Carry one product story across LinkedIn, sales email, landing page, and video.
Every model returned the same bounded page structure. The shared renderer
controlled fonts and basic layout. The model still decided what the reader saw
first, which evidence belonged together, how much detail to reveal, and which
action followed.
What happened across all 42 rows
Model
Passes
Contract failures
Provider errors
Measured cost
Kimi K3
3
0
4
USD 0.1491480
GLM-5.2
7
0
0
USD 0.0781736
GPT-5.6 Sol
7
0
0
USD 0.6756450
Claude Fable 5
5
0
2
USD 0.8494000
Claude Opus 5
5
2
0
USD 0.8413650
Qwen3.7-Max
7
0
0
USD 0.3110175
Total
34
2
6
USD 2.9047491
Provider errors and content failures answer different questions. Kimi K3 hit
four Moonshot overload errors, while Claude Fable 5 hit two Anthropic overload
errors. Those rows tell us about service capacity during this run window, not
the quality of an artifact that was never returned. Kimi later passed three
tasks, so the evidence does not support calling the route generally unavailable.
Claude Opus 5 returned both paid contract failures. One onboarding answer ended
with malformed JSON. One campaign answer omitted required evidence fields. We
kept both failures and did not repair or retry them.
GLM-5.2, GPT-5.6 Sol, and Qwen3.7-Max passed all seven frozen contracts. That is
the best contract record in this one run. It is not an overall model ranking.
One fixture, one prompt wording, one generation, and one renderer cannot settle
clarity, taste, conversion, or long-run reliability.
The later Qwen3.8-Max extension passed six of the same seven decision contracts.
Its two-depths call timed out without a usable answer. That row is a service
outcome from this run window, not evidence that the model would fail the content
contract if an answer were returned.
Compare two decision-page tasks
The first six routes completed both rows. Qwen3.8-Max completed the accessibility repair, while its two-depths call timed out. Open any image to inspect the full page at a readable size.
Kimi K3
Executive and practitioner viewsAccessibility repair
GLM-5.2
Executive and practitioner viewsAccessibility repair
GPT-5.6 Sol
Executive and practitioner viewsAccessibility repair
Claude Fable 5
Executive and practitioner viewsAccessibility repair
Claude Opus 5
Executive and practitioner viewsAccessibility repair
Qwen3.7-Max
Executive and practitioner viewsAccessibility repair
Qwen3.8-Max
Accessibility repair
What the clean rows let us compare
The original six models passed the executive-and-practitioner task and the accessibility
repair task. That gives readers two complete rows without hiding missing models.
The screenshots still differ in headline, grouping, density, labels, and visual
emphasis even though every page met the same mechanical rules.
The useful next judgment is human: which page lets a reader recover the main
decision fastest, then inspect the evidence without getting lost? The automatic
checks cannot answer that question. They only ensure that the comparison begins
with valid, evidence-bounded work.
Scope: This run covers seven structured communication tasks for one
fictional product on one date. It does not establish general intelligence,
formal accessibility compliance, conversion performance, professional
replacement, or a stable frontier-model ranking.
The honest result from the first smoke
This is not a frontier-model leaderboard.
The strongest statement supported by the earlier five-model smoke is narrow:
all five routes
produced one complete, fact-constrained carousel. Four produced a video plan
within the frozen schema, while Opus 5 exceeded the outro limit. Four produced
a landing page within the frozen contract, while GLM-5.2 returned too few body
sections.
The smoke does not identify a winner. It does not establish universal
intelligence, conversion performance, professional-design replacement,
long-run reliability, native media understanding, or facility compliance.
It does establish a more useful way to compare models: keep the facts fixed,
show the actual work, preserve failures, separate the model from the renderer,
and ask readers what they understood before revealing the name.
Cost and reproducibility
The fifteen benchmark calls cost USD 0.4668652. Two earlier Opus 5 diagnostic calls
exposed field limits that the parser enforced but the prompt had not stated.
Those calls cost another USD 0.06343 and are disclosed as harness-development
spend, not benchmark results.
The larger frozen benchmark contains a LinkedIn post plus creation, editing,
adaptation, and repair tasks across four forms of communication. Three repeats
across 12 tasks and five models would produce 180 calls. The observed smoke
usage produced a 25 percent buffered estimate of USD 12.5721. That is a planning
estimate, not a guarantee.
We also packaged the fixtures, tasks, checks, renderers, stored answers, result
summaries, and dated run records in the public
Instavar visual communication benchmark
repository. Reproduction means replaying the same contract and renderer or
creating a new dated sample. Provider behaviour is not deterministic, so it
does not mean receiving the same words again.
What comes next
The next useful experiment is a blind reader test on the two complete
decision-work rows. Readers should first answer simple questions: what changed,
what is uncertain, and what should happen next? Only then should they rank
clarity and visual appeal or see the model names.
That test would move the comparison from contract compliance toward the thing a
business actually cares about: whether a person understood the evidence and
could make the intended decision.
None. The smoke used one generation per model per task and had no completed
blind reader review. Naming a winner would go beyond the evidence.
Why did every carousel and page share one style?
The shared renderer fixed typography, colours, spacing, image handling, and
responsive behaviour. That makes differences in message, order, and density
easier to inspect.
Why can readers see failed outputs?
A stored answer can be rendered without satisfying the task. GLM-5.2 returned
too few landing-page sections. Opus 5 exceeded the video outro limit. Both
remain failed even though readers can inspect them.
Are the lower video notes subtitles?
No. They describe what is visible during each period. Their position resembles
spoken subtitles, which is a known limitation of the renderer.
Did the models watch the raw video?
No. They planned edits from the same timestamped scene inventory. This task does
not measure direct video understanding.
Did the models build the production pages?
They wrote bounded page specifications. Instavar's shared renderer supplied the
React presentation layer, responsive behaviour, fonts, spacing, and image
handling.
Why use a fictional product?
A fictional product gives every model the same closed set of facts. No model
can benefit from hidden knowledge about a real brand, and invented claims are
easier to detect.