VariantBench Evaluates Genetics Agents

kenbwork · x · 2026-07-17

This post introduces VariantBench, a new benchmark designed to evaluate verifiable agent capabilities in variant discovery, statistical genetics, and personal genomics tasks.

It includes links to the paper, results page, and blog, noting that the project is led by @chriswzou. Results show that GPT-5.6 Sol / Codex and Claude Opus 4.8 Max / Pi hover around a 42.1% pass rate. With no configuration passing more than half of the attempts, the benchmark remains highly challenging for current agents.

Related event: Latch Releases VariantBench for Genomic Agents(3 posts)→

Original post →

More from coding & agent

coding & agent channel →