Simulated-persona blind test: Fable 5.1 vs Opus 5.5 over 12 held-out rounds
every · x · 2026-09-23
The author describes an interesting model comparison method: personas were tested by simulating 'dan' and having models engage in a 12-round discussion on the real positions he held out in real life, comparing Fable 5.1 vs Opus 5.5 in Kieran Klaassen's Cozy Island setup.
More from coding & agent
- OpenAI Launches Prompt Caching Dashboard to Track Cache-Hit Rates and Misses — OpenAIDevs · 2026-09-23
- Firecrawl Launches Alexandria, Raises $75M Series B to Catalog the Internet for Agents — devdigest · 2026-09-23
- Listeners, not speakers: new study flips agent communication with interruptible generation — lileics · 2026-09-23
- CMU × Meta's HANDRAISER Cuts Multi-Agent Communication Cost 32.2% by Learning to Interrupt — lileics · 2026-09-23
- Beware: letting coding agents fix TS errors with 'as any' quietly kills compiler feedback — gethackteam · 2026-09-23
- Reviewing AI coding plans is a waste of time — get 3 tested solutions instead — alex_frantic · 2026-09-23