Meta releases WearableQA, a benchmark testing LLM health reasoning on real wearable data
meta · hf · 2026-09-10
Meta has released WearableQA on Hugging Face, a benchmark for evaluating how well large language models reason over real-world wearable device data in health contexts.
- Questions are multiple-choice, derived from real longitudinal wearable data
- Evaluation spans both data dimensions (understanding time-series signals like heart rate and activity) and health dimensions (health-related reasoning)
It offers a new evaluation tool for models working with personal health data.
More from Research
- ReactVAU at ECCV: slow-fast decoupling cuts MLLM calls for streaming video anomaly understanding — CMHungSteven · 2026-09-10
- ReactVAU at ECCV 2026: fast-slow streaming video understanding cuts heavy MLLM calls — CMHungSteven · 2026-09-10
- Perplexity releases Q2D-Web: a 190M-document retrieval benchmark for agentic RAG — antoine_chaffin · 2026-09-10
- Correction: the SmolVLM data-filtering trick is from DeepSeek's earlier tech report — eliebakouch · 2026-09-10
- DeepSeek used SmolVLM to quality-filter interleaved pretraining data, researcher spots — eliebakouch · 2026-09-10
- SmolVLM used for strict image-text quality scoring in new model tech report — eliebakouch · 2026-09-10