AMA: 167 models charted on a creative-writing benchmark Pareto frontier
OnlyProggingForFun · reddit · 2026-09-27
A Redditor opened an AMA about their creative-writing benchmark, which now covers 167 models and plots a Pareto frontier of performance vs cost. They invite questions on specific models, costs, tasks and scores.
More from Models
- Ex-NVIDIA engineer: US labs ignored global users, so the world runs on Chinese models — ivan_bezdomny · 2026-09-27
- GPT-OSS chat template bug silently drops past answers, degrading multi-turn coherence — arbv · 2026-09-27
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27