Protein design preprint compresses sequences into 8,192 ProtWords and generates folded proteins
CatAstro_Piyush · x · 2026-07-25
- The post points to a preprint that compresses protein sequences into a latent vocabulary of 8,192 ProtWords, downsampled to L/4, while preserving both sequence and fold information.
- The authors say they can perform diffusion sampling in this space to generate folded proteins from scratch.
- The thread also mentions functional cofilin designs reported in the preprint, suggesting the method can produce not just plausible folds but experimentally meaningful designs.
More from Research
- Meta, NYU and Yann LeCun argue BERT-style encoders break under scaling — pbaylies · 2026-07-25
- Bayesian thinking left room in LLM sampling, but the field spent three years on structured generation — remilouf · 2026-07-25
- AIMACS26 talk spotlights structured LLM outputs and Lean CSLib — swarat · 2026-07-25
- Paper argues LLMs are “anthropomimetic,” mirroring human flaws as well as strengths — dioscuri · 2026-07-25
- An AI detector correctly labels four AI essays and two human ones, then the writer jokes that only machines can tell — paul_cal · 2026-07-25
- HartwigGroup joins Genesis Mission to build ML models for synthetic chemistry — CatAstro_Piyush · 2026-07-25