Stanford prof warns BioReason benchmark has task definition flaws and prompt leakage
anshulkundaje · x · 2026-10-03
Stanford professor Anshul Kundaje argues the so-called "multimodal biological reasoning" model isn't really that, and its BioReason benchmark has serious task-definition problems plus information leakage through prompts. He advises against using these tasks for benchmarking or building on those task definitions.
Related event: Stanford's Kundaje blasts AI×Bio rigor gaps and NeurIPS review quality(11 posts)→
More from Research
- Math PhD Trains AI to Crack Open Research Problems, Reports Progress in Arithmetic Physics — shuchaobi · 2026-10-03
- childes-db 2026.1 Released: Child Language Database Grows to 24.2M Utterances with New Annotations — najoungkim · 2026-10-03
- SETA Terminal Agent Training Suite Accepted at NeurIPS, Open-Sources 4,500+ Verifiable RL Environments — Thom_Wolf · 2026-10-03
- MedARC journal club: pretraining foundation models for intracranial EEG — iScienceLuvr · 2026-10-03
- Mosaic: exact constrained decoding for diffusion LLMs via finite automata, NeurIPS paper — StefanoErmon · 2026-10-03
- NYU Researchers Challenge Anthropic's Claim That LLMs Can Introspect — tallinzen · 2026-10-03