RareBench results: Gemini 3.7 Flash leaps ahead, DeepSeek Pro shows no gain
danielmckinn0n · x · 2026-08-18
Gamowlabs published a RareBench (rare-disease genomics benchmark) blog post, with paper and data coming soon. Key findings:
- Gemini 3.7 Flash is a huge leap over Flash 3.6. If Google DeepMind can flip the remaining yellow categories to green, it would be near SOTA while fast and reasonably cheap. The author adds he loves the model for interactive genomics research use-cases — smart enough in the domain that its speed really accelerates online hypothesis generation.
- DeepSeek Pro results confirmed: despite overwhelming sentiment that the model was a great overall leap forward, it shows no improvement over Flash on RareBench.
More from Models
- Grok 4.6 coding test: nails complex logic but systematically misses basics — mark_k · 2026-08-18
- Qwen3.8-27B Benchmarks Show It Neck and Neck with DeepSeek V4 and GPT-5.6 Luna Max — anderspitman · 2026-08-18
- Orion 16B hits 100B training tokens using DPP on distributed GPUs — markjeffrey · 2026-08-18
- DeepSeek V4-Pro goes GA with configurable reasoning effort and Responses API — thione · 2026-08-18
- xAI releases Grok 4.6, focused on long-running agents, matching GPT-5.6 Sol on AA index — thione · 2026-08-18
- Claude to add invisible watermarks to AI text under EU rules, changing how it picks words — nordicinst · 2026-08-18