VariantBench Evaluates Genetics Agents
kenbwork · x · 2026-07-17
This post introduces VariantBench, a new benchmark designed to evaluate verifiable agent capabilities in variant discovery, statistical genetics, and personal genomics tasks.
It includes links to the paper, results page, and blog, noting that the project is led by @chriswzou. Results show that GPT-5.6 Sol / Codex and Claude Opus 4.8 Max / Pi hover around a 42.1% pass rate. With no configuration passing more than half of the attempts, the benchmark remains highly challenging for current agents.
Related event: Latch Releases VariantBench for Genomic Agents(3 posts)→
More from coding & agent
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- LangSmith adds tracing for Pipecat, LiveKit, OpenAI Realtime, and Gemini Live — LangChain · 2026-07-22
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- Annotated transcript of a Claude Code team interview is now available — trq212 · 2026-07-22
- Claude Code skill uses 10 Markdown rules to make outputs ADHD-friendly — alex_verem · 2026-07-22