Eight frontier models. One product story.

Kimi K3, GLM-5.2, GPT-5.6 Sol, and five other model routes each explained the same fictional product as a carousel, video, and landing page. The finished work makes their different choices easy to see.

The short answer: no model won overall. Most outputs followed the brief, but they guided the eye, ordered the story, and handled sound in noticeably different ways. GLM also returned too little for one landing-page task. This is one run per task, not a universal ranking.

Carousel
Both passed. Different openings.
Video
Both passed. Kimi muted the clip; GLM kept its sound.
Landing page
Kimi passed. GLM returned too little.

Eight routes tested: Kimi K3, GLM-5.2, GPT-5.6 Sol, Claude Fable 5, Claude Opus 5, Qwen3.7-Max, DeepSeek V4 Flash, and Qwen3.8-Max. Not every route appears in every run.

How we kept the comparison fair
## Why this comparison is useful AI can produce more copy, pages, and media plans than most teams can review. The scarce resource is no longer the first draft. It is judgment. A marketer needs to know whether the offer is clear. A product manager needs a first-time visitor to understand what the product does and what to do next. A designer looks for hierarchy, rhythm, and clutter. An engineer needs to separate a model mistake from a rendering mistake. An operator or finance lead cares about the cost of a usable attempt, not simply the cheapest call. Most model comparisons compress those questions into one score. Communication does not have one correct answer, so we kept the questions separate. | Reader | The decision this comparison can inform | What it cannot settle | | ---------------------- | ------------------------------------------------------------- | --------------------------------------- | | Marketing | How the same offer can be framed and sequenced | Which version will convert | | Product | Whether the promise and next step are easy to find | Whether real customers want the product | | Design | Hierarchy, density, and eye movement within one visual system | Unrestricted art direction | | Engineering | Constraint following, failure handling, and rendering | General coding ability | | Operations and finance | Cost for these attempts and a bounded rerun process | Stable future pricing | **Scope:** These are useful questions for many reader-facing products. The evidence in this article comes from one fictional product, one lab walkthrough, one generation per model per task, and eight dated API routes. ## The test in plain language We created Morrowboard, a fictional browser tool for small launch teams. Its promise was deliberately simple: > Turn scattered launch notes into one decision-ready brief. Every model received the same approved facts: * users can paste text or import Markdown files * Morrowboard sorts items into facts, decisions, and open questions * every item keeps a link to its source note * the finished brief can be exported to PDF or Markdown * a sample workspace opens without an account * the exact call to action is “Open the sample workspace” The brief prohibited invented customer counts, time-saving percentages, security guarantees, and unlisted integrations. The models returned structured plans rather than finished production code. One shared renderer supplied the fonts, colours, spacing, photography, responsive behaviour, and video export. Holding that layer constant makes differences in message, order, density, and editing choices easier to inspect. This design also creates a clear boundary. The comparison tests how the models shape information inside a visual format. It does not test native image generation, direct video understanding, unrestricted graphic design, or the ability to build an entire production website alone. ## How to read a “pass” A pass means that one response followed the frozen factual and structural rules. It does not mean that the artifact is beautiful, persuasive, easy to understand, or likely to convert. That distinction matters because a polished artifact can invent a statistic, while a valid artifact can still feel crowded or dull. The benchmark records contract compliance, rendering, cost, and reader preference as different questions.