Grok 4.6 Biology Eval: Matches Opus 5 Accuracy at a Substantially Lower Cost

kenbwork · x · 2026-08-14

Researchers evaluated the newly released Grok 4.6 model on short-horizon biology benchmarks comprising 1,716 trajectories.

The results indicate that Grok 4.6 achieves accuracy roughly on par with Opus 5 / GPT-5.6-Sol, while being substantially cheaper to run.

Original post →

More from Models

Models channel →