AI Drug Discovery Benchmark Update: General LLMs Closing Gap on Specialized Models

DeryaTR_ · x · 2026-08-03

Insilico Medicine updated its Drug Discovery and Development (DDD) benchmark portal, introducing comparisons between frontier general-purpose LLMs and specialized models.

On the multi-step synthesis planning benchmark URSAbench, general-purpose models are rapidly closing the gap. GPT 5.6 Sol (High) emerged as the top-ranked LLM, taking 3rd place overall, behind only two dedicated retrosynthesis planners.

However, specialized models still set the pace: RetroChimera by Microsoft broke the 30-point barrier to hold the #1 spot, followed by AstraZeneca's AiZynthFinder at #2.

Original post →

More from Research

Research channel →