Protein modeling thread says the real question is small MSA vs. big model
sokrypton · x · 2026-07-22
The post argues that the real comparison is small MSA vs. big model in protein modeling.
It says the biggest gains over pLMs appear in cases where there are very few sequences available. For large protein families, the practical baseline is often simply downloading related MSAs from AFDB and stacking your sequence on top.
The attached table highlights examples where MSA Pairformer / ESMC shows the largest ΔP@L improvements alongside the corresponding MSA depth values, suggesting that sequence context availability strongly shapes where the method helps most.
Related event: Protein Modeling: Small MSA vs Large Language Models(2 posts)→
More from Research
- Recursive self-improvement in AI shifts from bounded refinement to autonomous research loops — theomitsa · 2026-07-22
- A reposted AGI architecture blueprint points to system design for future general intelligence — theomitsa · 2026-07-22
- ASR paper boosts German disfluency F1 from 10% to 79% with verbatim control — nyralabs · 2026-07-22
- Kimi K3 shows benchmark awareness in 61% of trajectories, study says — gleech · 2026-07-22
- GPT models may be too obedient, raising paperclip-style alignment risks — aiamblichus · 2026-07-22
- AAAI 2027 submission count is joked to be around 50K in an academia meme — prajdabre · 2026-07-22