BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings
Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora, Keshav Balaji, Divyansha Lachi, Nanda H. Krishna, Jingyun Xiao, Yizi Zhang, Ximeng Mao, Wenrui Ma, Han Yu, International Brain Laboratory, Daniel Birman, Niccolò Bonacchi, Gaelle A. Chapuis, Joana A. Catarino, Felicia Davatolhagh, Mayo Faulkner, Laura Freitas-Silva, Fei Hu, Julia M. Huntenburg, Anup Khanal, Inês Laranjeira, Petrina Lau, Guido T. Meijer, Nathaniel J. Miska, Jean-Paul Noel, Alejandro Pan-Vazquez, Georg Raiser, Cyrille Rossant, Karolina Z. Socha, Anne E. Urai, Miles J. Wells, Steven J. West, Olivier Winter, Blake Richards, Guillaume Lajoie, Cole Hurwitz, Mehdi Azabou, Matthew R. Whiteway, Liam Paninski, Eva L. Dyer
cs.LG, q-bio.NC
2026-09-19
BrainWideBench tests across-animal transfer on IBL recordings from 139 mice and 276 regions. Pretraining beats matched single-session baselines, but no method wins all three suites of behavior, dynamics, and anatomy.
Dense electrophysiology can now record single neurons across many animals and many brain areas. Analysis still mostly fits one session or one region at a time. A general-purpose model of the mouse brain needs a shared split, a shared set of downstream jobs, and a way to score transfer to unseen animals.
Existing spike-resolution benchmarks either cover few animals or a single predictive task. BrainWideBench asks a stricter question: pretrain on many mice, then reuse one set of weights on held-out mice for behavior decoding, neural activity prediction, and anatomical identification.
The source is the International Brain Laboratory Brainwide Map: 139 mice on a visually guided decision task, Neuropixels coverage of 276 regions, more than 600 hours raw. Pretraining uses 126 mice and 423 sessions. Evaluation holds out 13 mice, 29 sessions, and 124 regions. Quality-control metadata is released at session, probe, and unit level.
Three suites aim at different properties.
Baselines are grouped as task-suite-supervised (TSS) versus unsupervised relative to the suite (TSU), plus heavily tuned single-session models.
The headline is already in the abstract: pretraining beats matched single-session baselines, but gains track the match between pretraining objective and downstream suite. No method is uniformly best.
On TS1, POYO+ (pretrained on decoding) has mean rank 2.53 and POSSM 3.18. Single-session POYO moves from 4.80 to 2.53 as POYO+; NDT moves from 6.80 to 3.58 as NDT-Stitch. Lick D² is 0.678 for POYO+ versus 0.151 linear and 0.549 for a single-session CNN. Choice balanced accuracy is 0.689 versus 0.614 linear. Self-supervised NDT-Stitch and MtM still transfer to whisker, right paw, choice, and reward, so the features are not locked to the original head. The overall ranking still favors behavior-supervised pretraining.
On TS2, pretrained models beat single-session latent-variable fits. Co-smoothing D²: MtM 0.191, NDT-Stitch 0.157, LFADS 0.176. Forecasting D²: NDT-Stitch 0.177, NDT 0.158, MtM 0.131. Spatial masking helps co-smoothing; temporal reconstruction helps forecasting, even when pretraining was non-causal.
TS3 opens a wider gap. Identity-centric NuCLR reaches 0.654 macro-F1 with a multi-unit linear probe; NEMO reaches 0.605 with a multi-unit MLP. POYO+ in the same protocol sits between 0.072 and 0.132. LOLCAT, trained directly on region labels, hits 0.473 multi-unit F1 and still trails NuCLR. Embeddings built for behavior decoding do not spontaneously sort by anatomy.
This is one of the few unified evaluations of across-animal transfer on multi-region single-cell data. For anyone training neural foundation models, “general-purpose” splits into three separately losable bets: behavior, dynamics, anatomy. Current methods are mostly designed for one of them.
The engineering moral is blunt. Single-session models need heavy hyperparameter search to peak. Pretrained models match or beat them with an almost standard fine-tune, amortizing that search. The remaining hole is also clear: behavior and dynamics models still grow unit embeddings or stitchers on the target session. Calibration-free transfer is still open, and this benchmark is built to score it.
The authors are explicit. One species, one modality, one behavioral paradigm. The 13-mouse eval set keeps compute and anatomical coverage in balance; coarse groupings are stable, within-group ranks move when different mice are held out, and should not be read as pairwise trophies. Labels such as TSS versus TSU are coarse baskets that do not capture each model’s inductive bias.
TS1 sessions last about 1.4 hours on average, so probe drift and spike-sorting quality are part of the test by design. TS2 dropped causal splits, which is an admission that unit identity is unstable over long timescales. TS3 uses 10 coarse regions, not cell types. If a paper claims a generalist, check which column it actually won.