Introducing the VariantBench Benchmark
kenbwork · x · 2026-07-17
The author has released VariantBench, a verifiable benchmark for variant discovery, statistical genetics, and personal genomics.
The post shares some initial results:
- GPT-5.6 Sol / Codex and Claude Opus 4.8 Max / Pi are leading the pack;
- The highest pass rate is approximately 42.1%;
- No configuration managed to pass more than half of all attempts.
This indicates that the benchmark is still exceptionally difficult for models, making it an excellent tool for assessing capabilities in tasks that combine bioinformatics with reasoning.
Related event: Latch Releases VariantBench for Genomic Agents(3 posts)→
More from Research
- Anthropic masterclass spotlights how to build and observe AI agents — _jaydeepkarale · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21