Protein modeling thread says the real question is small MSA vs. big model
sokrypton · x · 2026-07-22
The post argues that the real comparison is small MSA vs. big model in protein modeling.
It says the biggest gains over pLMs appear in cases where there are very few sequences available. For large protein families, the practical baseline is often simply downloading related MSAs from AFDB and stacking your sequence on top.
The attached table highlights examples where MSA Pairformer / ESMC shows the largest ΔP@L improvements alongside the corresponding MSA depth values, suggesting that sequence context availability strongly shapes where the method helps most.
Related event: Protein Modeling: Small MSA vs Large Language Models(2 posts)→
More from Research
- Researcher bootstraps from fly connectome to build increasingly intelligent connectomes — airkatakana · 2026-09-11
- CellFluxRL: RL-based biological grounding for virtual cell models, submitted to ECCV 2026 — Prof_Lundberg · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11