BehaviorBench: benchmarking frontier models on human behavior across 20 scenarios

Scobleizer · x · 2026-10-09

Researchers from Michigan, Stanford and others released BehaviorBench, a benchmark evaluating foundation models' understanding of human behavior across 20 scenarios and four capabilities, alongside a behavioral-science model family BeFM1.5 (4B and 70B).\n\n- Evaluation runs at both the individual and distributional levels, with a leaderboard offering Mean Win Rate / ELO rankings\n- Reasoning models are tagged with the reasoningeffort used\n- Paper, dataset, code, and a BeFM chat endpoint are all public

Original post →

More from Research

Research channel →