1B Parameters Suffice to Match 70B Baseline in Predicting Human Cognition
socius · hf · 2026-08-10
While fine-tuned LLMs serve as general-purpose cognitive proxies, their scale requirements and processing mechanisms remain unclear. Researchers trained 14 models ranging from 135M to 14B parameters on Psych-101, a dataset of 10.7 million trial-level choices.
Experiments reveal that in-distribution, scale barely matters: models fall within a narrow band, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. However, a markedly steeper scaling gradient emerges out-of-distribution, with larger models clearly advantaged in generalizing to novel task structures.
By progressively stripping prompt channels, researchers found that masking stimuli and feedback content destroys 75.7% of learned information, pushing models below chance. This demonstrates that small cognitively fine-tuned models show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by training paradigms.
More from Research
- Stanford's ChatEHR Deployment: $6M+ Estimated First-Year ROI and Inadequacies of Benchmarks — EricTopol · 2026-08-10
- ICML 2026 Machine Unlearning Tutorial Released with Videos and Slides — thegautamkamath · 2026-08-10
- SceneGen Generates 3D Scenes from a Single Image in One Feedforward Pass — tom_doerr · 2026-08-10
- Surya Ganguli Shares Top Summer Schools in Computational Neuroscience and AI — SuryaGanguli · 2026-08-10
- IFDS Research Expo Aug 10-12 at UW-Madison: AI foundations, ML, stats, optimization — prof_kamilov · 2026-08-10
- Primus Launches Autonomous ML Research Agent: 30x Faster from Hypothesis to Paper — JayAlammar · 2026-08-10