Frontier models barely improve at medical synthesis, best F1 just 41.1
manoelribeiro · x · 2026-10-07
Frontier models are not improving much at medical conclusion synthesis: the best performer, GPT-6.1 Sol, reaches only 41.1 factual F1, with most models clustered around 0.30–0.33, showing healthcare synthesis remains a hard task.
More from Models
- Ollama hosts Google's EmbeddingGemma 2, a 740M multimodal embedding model for on-device use — ollama · 2026-10-07
- Reddit users grow frustrated with ChatGPT's over-refusals on innocuous prompts — Crixusgannicus · 2026-10-07
- Runware launches two API content moderation models that take plain-language policies — aziz4ai · 2026-10-07
- Ethan Mollick to AI Labs: Make sure your models actually understand your own products — emollick · 2026-10-07
- OpenAI launches Decisions API in public beta, up to 10x faster than GPT-6 Luna — OpenAIDevs · 2026-10-07
- Cloudflare's open-source vision decision model clef impresses developers — michellechen · 2026-10-07