Blind test of 90 page pairs: named style references beat adjectives 68% across Opus, Sol, Kimi
maxsloef · x · 2026-09-27
maxsloef built a small style-elicitation eval: Opus 5.5, Sol 6, and Kimi K3 each generated two landing pages for a fake startup — one from an adjective set (dreamy, gentle, luminous), one from a named style reference (e.g., James Turrell) — and he blind-rated 90 pairs.
- Named references won 68% of the time: Sol 73%, Kimi 67%, Opus 63%.
- The effect holds even when the model picks the reference itself.
- He released the eval as Namedrop and hopes labs will hillclimb on it.
Related event: Named reference beats adjectives for style prompting in blind test(2 posts)→
More from Models
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma message in two days — luisdans · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27