Frontier models, tested on work you can inspect.

We gave dated GPT, Claude, Qwen, GLM, Kimi, DeepSeek, and Tencent routes the same bounded work. You can inspect the finished visual artifacts, the code checks, the repeated agent outcomes, and the cost of each route.

The short answer: no model won every kind of work. Qwen3.8-Max-0902 led the small 31-check screen. Six hosted profiles preserved the important external behaviour in all 16 repeated agent runs. GLM-5.3-Flash was the least expensive of those six. Harder tasks still exposed incomplete work, unsafe attempts, and state-recovery failures.

Small screen
Qwen3.8-Max-0902 led with 31/31 after audit.
Repeated work
Six hosted profiles preserved all 16 important outcomes.
Harder work
No route completed all three longer agent tasks.

The article combines five dated stages. Earlier visual-only routes remain visible, while the matched agent tables cover the model profiles that ran the same code, evidence, state, and safety contracts. Not every route ran every stage, so missing cells are not treated as failures.

How we kept the comparison fair
## Why this comparison is useful AI can produce more copy, pages, and media plans than most teams can review. The scarce resource is no longer the first draft. It is judgment. A marketer needs to know whether the offer is clear. A product manager needs a first-time visitor to understand what the product does and what to do next. A designer looks for hierarchy, rhythm, and clutter. An engineer needs to separate a model mistake from a rendering mistake. An operator or finance lead cares about the cost of a usable attempt, not simply the cheapest call. Most model comparisons compress those questions into one score. Communication does not have one correct answer, so we kept the questions separate. | Reader | The decision this comparison can inform | What it cannot settle | | ---------------------- | ------------------------------------------------------------- | --------------------------------------- | | Marketing | How the same offer can be framed and sequenced | Which version will convert | | Product | Whether the promise and next step are easy to find | Whether real customers want the product | | Design | Hierarchy, density, and eye movement within one visual system | Unrestricted art direction | | Engineering | Constraint following, failure handling, and rendering | General coding ability | | Operations and finance | Cost for these attempts and a bounded rerun process | Stable future pricing | **Scope:** These are useful questions for many reader-facing products. The visual evidence comes from one fictional product, one lab walkthrough, and one generation per model per task. The later agent evidence uses small controlled fixtures and at most two repetitions per task. Neither part establishes a general model ranking. ## The test in plain language We created Morrowboard, a fictional browser tool for small launch teams. Its promise was deliberately simple: > Turn scattered launch notes into one decision-ready brief. Every model received the same approved facts: * users can paste text or import Markdown files * Morrowboard sorts items into facts, decisions, and open questions * every item keeps a link to its source note * the finished brief can be exported to PDF or Markdown * a sample workspace opens without an account * the exact call to action is “Open the sample workspace” The brief prohibited invented customer counts, time-saving percentages, security guarantees, and unlisted integrations. The models returned structured plans rather than finished production code. One shared renderer supplied the fonts, colours, spacing, photography, responsive behaviour, and video export. Holding that layer constant makes differences in message, order, density, and editing choices easier to inspect. This design also creates a clear boundary. The comparison tests how the models shape information inside a visual format. It does not test native image generation, direct video understanding, unrestricted graphic design, or the ability to build an entire production website alone. ## How to read a “pass” A pass means that one response followed the frozen factual and structural rules. It does not mean that the artifact is beautiful, persuasive, easy to understand, or likely to convert. That distinction matters because a polished artifact can invent a statistic, while a valid artifact can still feel crowded or dull. The benchmark records contract compliance, rendering, cost, and reader preference as different questions.