Meta Researchers Release WearableQA, a 4,084-Question Benchmark for LLM Health Reasoning
iScienceLuvr · x · 2026-09-08
Meta researchers introduce WearableQA, a benchmark for evaluating LLMs on reasoning over longitudinal wearable data.
- Built from 200 real users' wearable records: hundreds of days per person across 16 daily metrics, plus a 17-biomarker blood panel and demographic info
- 4,084 multiple-choice questions total
- Two axes: data reasoning (trends, anomalies, relationships from raw measurements) and health reasoning (clinical interpretation)
The work addresses the lack of systematic evaluation as LLMs are increasingly used to analyze longitudinal wearable data for personalized medicine.
More from Research
- Training a 441M-Param Text-to-Image Diffusion Model From Scratch on a Single Local GPU — ostrisai · 2026-09-08
- 417k-param RNN generates all 6,573 frames of Bad Apple from a single initial state — SEBADA321 · 2026-09-08
- Can current LLM architecture reach AGI? An engineer lays out his doubts — mostly_deterministic · 2026-09-08
- MaintainabilityBench: grade AI on the cost of adding features, not correctness — kuza55 · 2026-09-08
- No training needed: strong model's harness lifts GPT-5.4-mini from 0.488 to 0.912 — jiqizhixin · 2026-09-08
- A prerequisite-free elementary introduction to information geometry — FrnkNlsn · 2026-09-08