1B Parameters Suffice to Match 70B Baseline in Predicting Human Cognition

socius · hf · 2026-08-10

While fine-tuned LLMs serve as general-purpose cognitive proxies, their scale requirements and processing mechanisms remain unclear. Researchers trained 14 models ranging from 135M to 14B parameters on Psych-101, a dataset of 10.7 million trial-level choices.

Experiments reveal that in-distribution, scale barely matters: models fall within a narrow band, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. However, a markedly steeper scaling gradient emerges out-of-distribution, with larger models clearly advantaged in generalizing to novel task structures.

By progressively stripping prompt channels, researchers found that masking stimuli and feedback content destroys 75.7% of learned information, pushing models below chance. This demonstrates that small cognitively fine-tuned models show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by training paradigms.

Original post →

More from Research

Research channel →