PhenoBench: 90 tasks testing how well LLMs exploit deep phenotyping data
segal_eran · x · 2026-09-17
Built on the Human Phenotype Project, PhenoBench offers 90 tasks across 15 clinical domains using clinical, imaging, molecular, and wearable data to test what today's AI models can extract from a deeply phenotyped human cohort.
Key findings:
- Tabular foundation models performed best overall, but gains over ridge regression were small
- LLMs showed a classic "jagged frontier": strong on some tasks, weak on others
- Models trained on the same input fields generally did better
The authors note most effort went into defining tasks and grading answers—far more than actually running models—after which trying new models or measurements became cheap.
More from Research
- Cisco's 3B Model Matches GPT-5.5 on Bug Localization (0.223) With Far Fewer False Positives — shashib · 2026-09-17
- Parallel Decoding Distillation pushes generation efficiency for autonomous driving world models — abursuc · 2026-09-17
- NVIDIA's SpatialClaw uses code as action interface, beats prior agent by 11.2 points on 20 benchmarks — CMHungSteven · 2026-09-17
- World model interactivity: handling mixed-frequency state updates at 10Hz — abursuc · 2026-09-17
- Researcher: AI math 'darlings' long relied on fake baselines, math lacks empirical tradition — RexDouglass · 2026-09-17
- CoLLAs 2026 keynote: Is continual learning trapped in obsolete abstractions? — apsarathchandar · 2026-09-17