Benchmarking Frontier Models on Creative Tasks
soleio · x · 2026-07-11
Contra Labs aims to provide companies with the "most interesting and actionable" research. The quote highlights their head-to-head test of 4 frontier models: working creatives conducted blind evaluations based on 10 identical landing page briefs.
The conclusions are:
- GPT 5.6 Sol performed best under "loose, ambiguous briefs," with a win rate of 82%.
- However, when the task shifted to "explicit design specs," it ranked dead last.
- The author emphasizes that the same model's performance varies drastically under different constraints, meaning creative workflows rely heavily on how specific the input is.
Related event: Frontier Models Face Off in AI Web Design Benchmark(2 posts)→
More from Apps
- Tip: keep separate Chrome, Cursor, ChatGPT, Claude and Notion profiles — msg · 2026-07-21
- AI-made 3D website turns the homepage into a three-tower mini game — techartist_ · 2026-07-21
- Tesla expands Grok in-car assistant to five Asian markets — XFreeze · 2026-07-21
- Tesla’s FSD v14 Lite is reportedly headed to 4 million older HW3 cars — MatthewBerman · 2026-07-21
- DevFest Lisbon keynote will cover Google AI Studio’s latest vibe coding and agentic AI features — gerardsans · 2026-07-21
- Google says AI Search is still sending billions of clicks to websites each week — SEO · 2026-07-21