MSA can still beat big models when protein families have very few sequences
sokrypton · x · 2026-07-22
MSA will still beat bigModel in the hardest cases, according to the thread:
- The key advantage appears when a protein family has very few sequences. In that regime, multiple sequence alignments can still provide signal that a pLM struggles to infer alone.
- For large protein families, the author argues you can often just pull related MSAs from AFDB and place the query sequence on top.
- The comparison is framed as smallMSA vs bigModel rather than a blanket “MSA vs model” debate.
- A table credited to @yoakiyama is cited as supporting evidence.
The broader point: alignment-based priors still matter a lot when data are scarce, even as large protein language models improve.
Related event: Protein Modeling: Small MSA vs Large Language Models(2 posts)→
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11