pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
Tong Chen, Maximilian Holsman, Lin Zhao, Pranam Chatterjee
NeurIPS 2026
cs.LG, q-bio.BM
2026-10-01
Doob-h guided edit flows shrink biomolecules under hard constraints: GFP down 10 residues still glows in cells; Cas9 edits keep 100% PAM match vs 35%/0% for baselines.
Therapeutic biomolecules are rarely designed from scratch. The routine path is to edit an existing sequence: change a few positions, delete dispensable segments, trade potency against manufacturability and delivery size. Cas9 is the sharpest case; at 1368 residues, SpCas9 overflows AAV vectors, so in vivo delivery stays bottlenecked by size and shrinkage is a standing need.
Two models target shrinkage directly: RayGun encodes variable-length sequences into a fixed-dimensional space and decodes shorter ones, and SCISOR learns deletion planning with discrete diffusion. Neither jointly optimizes multiple developability objectives nor enforces hard constraints. Evolutionary search and RL can approximate Pareto fronts but need repeated online oracle calls. Offline, multi-objective, hard-constrained editing with insertions and deletions on variable-length sequences had no method covering all four at once.
The backbone is a pretrained Edit Flow, a continuous-time Markov chain over variable-length sequences whose steps are single insertions, deletions, or substitutions. pCoMole never retrains it; it only reshapes transition rates at sampling time, in three moves.
Constraints are checked only at termination, and intermediate states may leave the feasible set. Forcing feasibility at every step would block coordinated edit paths that break first and rebuild later. The relaxation is deliberate.
GFP shrinkage from the 238-residue wild type to a fixed 213 residues, averaged over 100 candidates per method:
| Method | Excitation error (nm, ↓) | Predicted brightness |
| RayGun | 22.36 | 9.42 |
| SCISOR | 20.13 | 19.03 |
| pCoMole (generic UniRef backbone) | 5.82 | 43.51 |
| pCoMole (GFP-tuned backbone) | 8.44 | 49.01 |
Even the backbone that never saw GFP gets the best excitation alignment, so the gain comes from guided sampling rather than a domain prior. Under a one-hour budget, only pCoMole reaches the strictest tier (3× wild-type brightness, ≤5 nm error), in 665 s; SCISOR is faster at the loose tier (14 s vs 266 s) and the ordering flips as the tier tightens.
The wet lab is the strongest evidence. Starting from a 239-residue in-house eGFP, two 229-residue designs (10 deletions each, one and two substitutions) retained clear green fluorescence under 488-nm imaging in BL21 cells, while a manually picked 228-residue control (11 deletions, 2 substitutions) showed no comparable signal.
Cas9: the PAM (the short protospacer-adjacent motif a Cas9 needs to recognize a target) determines which sites can be edited. Four orthologs shrunk by 52 to 157 residues with 100% PAM match and Cas9-likelihood held at 0.95 or above. Re-scoring with CICERO, an evaluator never used during guidance, exact PAM match rose from 4% to 88% for St3Cas9 and from 20% to 72% for GeoCas9. On a fixed-length benchmark (984→934 residues), pCoMole scored 100% PAM match against SCISOR's 35% and RayGun's 0%; time-to-first-usable stayed nearly flat across quality tiers (132/129/139 s) while SCISOR climbed to 240 s at the top tier.
Peptidomimetics: seven drug properties optimized jointly. Semaglutide-derived designs compressed from 248 to 186 tokens with predicted half-life up from 1.86 h to 4.03 h and affinity essentially unchanged (0.947→0.937); AutoDock VINA confirmed the designs still occupy the original binding pocket. An ablation gave augmented Tchebycheff a Pareto coverage of 0.73 versus 0.65 for plain Tchebycheff and 0.59 for weighted sums.
Controllable discrete generation has largely stopped at single objectives and soft penalties. pCoMole is a general recipe: any pretrained edit model, plus tunable weights and a feasibility gate, becomes a constrained Pareto sampler with error bounds on the approximation. Wet-lab confirmation is rare in generative protein design, and two designs both glowing in cells beats another metrics table. The cost is compute; rollout guidance is far more expensive than unguided sampling, which fits the drug-design regime where oracles are costly and every candidate counts.
The authors concede the compute overhead: 50-step sampling took 531 s per sequence in one peptidomimetic ablation. Wet-lab coverage is the bigger gap. Only GFP fluorescence was tested, and only qualitatively; the 3× brightness figure is an FPredX prediction with no matching measurement, and the 239-residue construct tested in the lab is not the 238-residue sequence used in the computational benchmark. Cas9 and peptidomimetic results rest entirely on predictors such as Protein2PAM and PeptiVerse, and guidance optimizes exactly those oracles, so their biases flow straight into the designs. The GFP comparison also forces RayGun and SCISOR into a fixed-length setting they were not built for, which is not entirely fair to the baselines; the Cas9 PAM gap stands on its own.