Zvi says Opus 5 matches Fable on virology tasks, raising a safety-policy question
TheZvi · x · 2026-07-25
Opus 5 appears competitive with Fable on virology tasks, Zvi says
In a post with benchmark charts, Zvi argues that Opus 5 looks at least as good at virology as Fable. His main takeaway is regulatory, not technical: if Opus 5 can perform at that level without Fable’s bio classifiers, then why does Fable need them?
The attached charts compare models on several virus-related tasks:
- AAV packaging rate classification on an unsupervised setup
- Long-form virology tasks for sequence, protocol, and end-to-end work
- Multimodal virology and DNA synthesis screening
Across the plots, Opus 5 is shown alongside Opus 4.7/4.8, Sonnet 5, and Mytho 5, with Opus 5 often near the top on the presented metrics. The post uses those results to question why one model is treated as more constrained than another if the capability gap is already so small.
Related event: Claude Opus 5 Virology Capabilities Spark Safety Debate(2 posts)→
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11