MSA can still beat big models when protein families have very few sequences
sokrypton · x · 2026-07-22
MSA will still beat bigModel in the hardest cases, according to the thread:
- The key advantage appears when a protein family has very few sequences. In that regime, multiple sequence alignments can still provide signal that a pLM struggles to infer alone.
- For large protein families, the author argues you can often just pull related MSAs from AFDB and place the query sequence on top.
- The comparison is framed as smallMSA vs bigModel rather than a blanket “MSA vs model” debate.
- A table credited to @yoakiyama is cited as supporting evidence.
The broader point: alignment-based priors still matter a lot when data are scarce, even as large protein language models improve.
Related event: Protein Modeling: Small MSA vs Large Language Models(2 posts)→
More from Research
- Protein language models can learn homo-oligomer contacts from single sequences — anshulkundaje · 2026-07-22
- Commercial frontier models blocked attack forensics because they misread the responder — morqon · 2026-07-22
- LeCun’s JEPA pitch gets a concrete world-model paper behind it — nikola_mr64990 · 2026-07-22
- MIT opens a large free AI library with classic books, lecture notes, and courses — ZabihullahAtal · 2026-07-22
- Research thread points to a finding from protein language models trained on sequence data — anshulkundaje · 2026-07-22
- Podcast Explores If a Single Cell Can Learn and Exhibit Cellular Memory — arjunrajlab · 2026-07-22