Anthropic's Fable Breaks RareBench Record After Relaxing Filters
danielmckinn0n · x · 2026-08-21
Anthropic's Fable model previously scored 0% on RareBench due to bio-related refusals. After relaxing filters, Fable took the lead with a staggering 41% + 7%, significantly beating @SpaceXAI @grok 4.6's previous record of 4.6%. The author plans to update the system to include Fable and rerun unsolved clinical genomes, expressing cautious optimism about diagnosing new cases.
More from Models
- Musk Confirms Work to Improve Grok's Writing Skills — mark_k · 2026-08-21
- Why 'Full Pass Rate' is a flawed metric for LLM evaluation — xeophon · 2026-08-21
- ARC Prize Adds Model Comparison, Gemini 3.7 Flash Scores High — mhmazur · 2026-08-21
- NVIDIA Explains Omni-Models: Unified Architecture for Text, Images, Audio, Video, and Actions — NVIDIA Developer · 2026-08-21
- Monitors Detect Significant Behavior Shift in Claude Opus — altryne · 2026-08-21
- Users report GPT-4.1 Sol model suddenly became dumb with irrelevant answers — M-M103 · 2026-08-21