Brain signals guide LLM reasoning: NMI study boosts accuracy by up to 13.2 points
机器之心 · wechat · 2026-09-26
Researchers from Peking University, Tsinghua and Microsoft Research Asia published a Nature Machine Intelligence study probing how LLM internal representations align with human brain activity during deductive reasoning — and turning that alignment into functional gains.
Key findings:
- Testing 10 open models (Qwen, Llama, Mistral, Phi, Gemma; 1.5B–72B) against fMRI data from syllogistic and transitive reasoning tasks, model representations explained 76% of explainable variance in reasoning-related brain regions — with notably higher predictivity in reasoning/multiple-demand networks than the core language network. Granularity drops to 27% per reasoning type, indicating only partial alignment.
- NARI (inference-time intervention): a learned mapping to fMRI space yields neural-guided directions that steer hidden states; it achieved 100% error-correction coverage on items humans solved and models failed, with aggregated directions transferring across models and problems.
- NARF (fine-tuning): internalizing the guidance into parameters improved robustness across premise counts, orderings and novel logic types, and combined with standard language supervision added 2.2 points on average (up to 13.2) including transfer to propositional and first-order logic tasks.
The work shifts the question from "are models brain-like?" to "can brain signals improve models?", though limitations remain around fMRI temporal resolution and the narrow reasoning domain.
More from Models
- Puppy Kill Bench: most models refuse, GPT6-Luna just executes the kill tool — MetroidsSuffering · 2026-09-27
- Ethan Mollick: Opus 4.7-5 lost the 'Claude feel', Opus 5.5 brings it back — emollick · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Hands-on: Opus 5.5 high beats GPT-6 astra xhigh on real Pagespeed optimization — mazzaTalk · 2026-09-27
- Supersonic Labs open-sources Julia 1, a 144M-parameter CPU-runnable decision model — ThePrimeClock · 2026-09-27
- Asked Grok to teach Japanese kanji, it started inventing its own — JoeJustice · 2026-09-27