Protein modeling thread says the real question is small MSA vs. big model

sokrypton · x · 2026-07-22

The post argues that the real comparison is small MSA vs. big model in protein modeling.

It says the biggest gains over pLMs appear in cases where there are very few sequences available. For large protein families, the practical baseline is often simply downloading related MSAs from AFDB and stacking your sequence on top.

The attached table highlights examples where MSA Pairformer / ESMC shows the largest ΔP@L improvements alongside the corresponding MSA depth values, suggesting that sequence context availability strongly shapes where the method helps most.

Related event: Protein Modeling: Small MSA vs Large Language Models(2 posts)→

Original post →

More from Research

Research channel →