PhenoBench: 90 tasks testing how well LLMs exploit deep phenotyping data

segal_eran · x · 2026-09-17

Built on the Human Phenotype Project, PhenoBench offers 90 tasks across 15 clinical domains using clinical, imaging, molecular, and wearable data to test what today's AI models can extract from a deeply phenotyped human cohort.

Key findings:

The authors note most effort went into defining tasks and grading answers—far more than actually running models—after which trying new models or measurements became cheap.

Original post →

More from Research

Research channel →