Baker lab's SAPP tests hundreds of designed proteins a day at about $24 each

2026-08-21

SAPP runs arrayed expression and SEC in 48 hours at ~$24/design, median 96% clonal purity on 929 reactions. DMX cuts gene cost ~5x and produced nM RSV binders (cb13).

What problem this solves

De novo protein design can emit large sequence libraries in a day. Wet lab still tests a few dozen designs over weeks. Pooled screens scale to thousands or millions, but they usually read one or two properties such as stability or binding, and they almost never say whether a given sequence expresses, stays soluble, is monodisperse, or which oligomer it forms. The missing piece in the design loop is cheap, repeatable, assay-agnostic arrayed production.

The Institute for Protein Design at the University of Washington answers with two protocols. SAPP takes gene fragments to milligram-scale purification and standard analytics. DMX then cuts gene-synthesis cost by starting from oligo pools. Nature's PDF resolved to a login page. The write-up below uses the matching bioRxiv preprint, the open supplementary methods, and the paper's Source Data file.

Method

SAPP skips the hard-to-parallelize clone-sequence-pick loop and standardizes every purification step, with Python driving ordinary instruments. Amino-acid sequences are reverse-translated and codon-optimized with DNA chisel, checked for synthesizability through the IDT API, split into Golden Gate pairs when too long, and given uniform adapters. Cloning runs as 1 µL Golden Gate on an Echo dispenser, with a 5 µL hand-pipette recipe as backup. Expression in E. coli, His-tag purification, and analytical SEC then report yield, dispersity, and oligomeric state. Wall clock is 48 hours, about 6 of them at the bench. Twenty-five compatible vectors are on Addgene.

Synthetic DNA is more than 80% of SAPP's bill, so DMX starts from oligo pools, adds barcodes, demultiplexes sequence-verified arrayed clones on nanopore reads, and feeds those clones into SAPP. The supplement is a full recipe book. Robotics is optional.

Results

From eBlock fragments, Supplementary Table 2 prices SAPP at $23.98 per design, $21 of that DNA, or $2,302 per 96-well plate. Source Data matches the preprint figure captions.

MetricNNumber
GGA clonal purity929median 96.1%, mean 92.4%
LC-MS mass error862median −131 Da (N-terminal Met cleavage in E. coli)
Soluble yield576median 0.36 mg, max 1.22 mg
Biological replicates96 pairsyield and main-peak retention reproduce

Two applications sit in the Source Data. Redesigned fluorescent proteins (muGFP, SYFP2) are tracked by SEC main-peak position before and after 1 hour at 95°C, a proxy for whether the design still folds and still glows. De novo binders to RSV F include cb13, with IC50 values of 4.6–9.9 nM; across 41 IC50s the median is 0.47 nM and the lowest is 14 pM. Cryo-EM of cb13 bound to RSV F refines to 4.63 Å overall, with secondary structure visible near antigenic site III. That binding data is consistent with the abstract's "potent neutralization"; per-point neutralization curves sit in the unread main figures.

At 1,000 designs, DMX costs about $4.75 each against $22.24 of SAPP DNA, close to the advertised five-fold drop. At 2,000 designs the unit cost is $3.38 and the saving factor is 6.59. Demultiplex quality: 93% of wells have complete barcodes, 70% carry the correct gene, 92% are one gene per well, 8% are mixed, 30% are incorrect. The preprint states the pipeline is already the institute default, used across dozens of projects and tens of thousands of designs.

Why it matters

The bottleneck in AI protein design is no longer another in silico score. It is whether design-express-characterize can run at hundreds per day, in a format that can train an active-learning loop. SAPP measures manufacturability, not one more binding assay. The recipe uses stock reagents and optional open-source automation, and the Addgene vectors are public. Labs with an Echo, or with patience for a 5 µL pipette protocol, can try it. Patent filings are already in: 63/464,881 for SAPP and 19/768,102 for DMX.

Limitations

Skipping per-clone sequencing is how the throughput appears, and how the risk appears. Mean clonal purity is 92% over 929 reactions, and the minimum is 0, so later functional reads can mix in the wrong molecule. The host is E. coli: Met cleavage, inclusion bodies, and eukaryotic modification are outside the protocol. A median 0.36 mg covers SEC and many binding assays; scale-up is not given the same throughput. DMX still yields 30% incorrect genes, so part of the DNA saving is spent on repeats. Source Data contains instrument output for fluorescent proteins and RSV binders; screening policy, hit rates, and the neutralization setup are not in the retrieved narrative, so IC50 should not be read as a neutralization titer. "Tens of thousands of designs" is a preprint claim and is not repeated in the journal abstract.

Terms

Source

What people are saying

All paper explainers