Protein folding post-training lifts LLM reasoning: +3.23pp on all 10 benchmarks
SJUT1 · hf · 2026-10-05
Can learning to fold proteins teach language models reusable reasoning skills? This paper answers with data.
- FoldingCorpus: a protein-derived QA dataset where each solved structure yields thousands of exactly checkable spatial/topological statements
- Fold2Reason: post-trains on two complementary signals — discrete structural answers via the language head and continuous 3D geometry decoded from the same shared representations
- On FoldBench, structure prediction scores reach 2.7–3.5x those of Qwen3.5-9B
- Generalization: gains on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning; macro-average accuracy rises from 45.09% to 48.33% (+3.23 pp)
- Controls trained on random, synthetic, or shuffled structures show substantially smaller or negative gains
The takeaway: non-linguistic, structure-dense scientific data is a practical source of post-training supervision for broad reasoning.
More from Research
- UVM Wins up to $38M to Build AI 'Digital Twins' for Critically Ill Patients, Could Cut ICU Stays 25% — HealthcareLdr · 2026-10-05
- SimuVerity Benchmark: Best Agent Scores Only 42.86 on Engineering-Grade Simulink Generation — Ruiqi Zhang · 2026-10-05
- MIT's Local Support Learning Fixes Catastrophic Forgetting in LLMs Without Old Data — MIT · 2026-10-05
- HelixWorld: Real-Time Audio-Visual World Model Runs at 24 FPS on a Single GPU — NoizAI · 2026-10-05
- BlockRank: sparse attention makes LLM in-context document ranking faster — dejanseo · 2026-10-05
- Judea Pearl: logic and causal discovery are the two pillars of Western science — yudapearl · 2026-10-05