Eval design note: models self-select style references in a separate call to isolate the effect
maxsloef · x · 2026-09-27
maxsloef explains the self-elicitation design in his Namedrop eval: each model picks a named reference evoking the target adjectives. Because choosing the reference during page creation would trigger longer chain-of-thought reasoning, the selection happens in a separate call — cleanly isolating the reference effect.
Related event: Blogger runs Namedrop eval in 1.5 hours of focused attention(4 posts)→
More from Models
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma message in two days — luisdans · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27