Frontier models barely improve at medical synthesis, best F1 just 41.1

manoelribeiro · x · 2026-10-07

Frontier models are not improving much at medical conclusion synthesis: the best performer, GPT-6.1 Sol, reaches only 41.1 factual F1, with most models clustered around 0.30–0.33, showing healthcare synthesis remains a hard task.

Related event: SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis(10 posts)→

Original post →

More from Models

Models channel →