BehaviorBench: benchmarking frontier models on human behavior across 20 scenarios
Scobleizer · x · 2026-10-09
Researchers from Michigan, Stanford and others released BehaviorBench, a benchmark evaluating foundation models' understanding of human behavior across 20 scenarios and four capabilities, alongside a behavioral-science model family BeFM1.5 (4B and 70B).\n\n- Evaluation runs at both the individual and distributional levels, with a leaderboard offering Mean Win Rate / ELO rankings\n- Reasoning models are tagged with the reasoningeffort used\n- Paper, dataset, code, and a BeFM chat endpoint are all public
More from Research
- Why AI doesn't actually read words: from BPE subwords to byte-level models like BLT and H-Net — jbhuang0604 · 2026-10-09
- NVIDIA details HSTU recommender inference stack with up to 5.93x lower latency — PyTorch · 2026-10-09
- Preprint: LLMs store numbers as curves and helices, but compute comparisons differently — tweetsatpreet · 2026-10-09
- Best model was cheapest: open-weights model ran 669 clinical decisions for 1.7 cents — antoine_chaffin · 2026-10-09
- Ofir Press Points to ExcelBench as the Right Scale for Benchmarking Coding Agents — OfirPress · 2026-10-09
- tangermeme, a genomic sequence-to-function modeling toolkit, published in Nature Methods — anshulkundaje · 2026-10-09