RLXF Feeds Wet-Lab Measurements Back into ESM-2 to Outdo Evolution on Protein Design
bravo_abad · x · 2026-09-13
Blalock and coauthors propose RLXF, using reinforcement learning to feed experimental measurements back into the protein language model ESM-2, steering generation away from "what evolution would choose" toward "what performs better in the lab":
- Case study: the light-sensing protein CreiLOV was never evolved to be bright. At position 43, ESM-2 prefers the natural cysteine, even though experiments showed swapping in alanine makes CreiLOV brighter.
- The mismatch between model preference and measured results is used as the RL signal to shift generation away from evolutionary bias.
- Notably, simply stacking the best individual mutations performed poorly, highlighting complex combination effects; the RL-guided sequence generation route matters more than finding a few good mutations.
More from Research
- Proxy Policy Steering adapts frozen VLA models to new tasks at inference time — weichiuma · 2026-09-14
- New research: standard SGD matches AdamW for LLM RL training, with far less memory overhead — zhaoran_wang · 2026-09-14
- RSI work separates practical harness self-improvement from unproven intelligence explosion — arthurcolle · 2026-09-14
- Amazon Proposes Query-Aware Index Pruning to Optimize Retrieval Under Budget Constraints — _reachsumit · 2026-09-14
- New Paper Finds Retrieval Signals Give No Reliable Routing Gain in Adaptive Multimodal RAG — _reachsumit · 2026-09-14
- Google: Graph RAG Cuts API Hallucination Rate from 56.4% to 16.2% in Java-to-Python Migration — _reachsumit · 2026-09-14