Blinded antibody contest: 29 labs, 511 designs, ranking models lose to random picks

2026-08-21

AIntibody tested 511 AI antibodies from 29 groups on SARS-CoV-2 RBD. Affinity maturation reached 94.7 pM, but cluster ranking mostly lost to random clone picking.

What problem this solves

Most computational antibody papers score models on public sequences and affinities after the fact. Without a blind wet-lab bake-off, it is hard to tell method progress from overfitting to released sets.

AIntibody runs a CASP-style prospective test. Twenty-nine organizations submitted 511 sequences. Organizers expressed them as full-length IgGs and scored them with high-throughput SPR, orthogonal KinExA, and a five-assay developability panel. The antigen is SARS-CoV-2 RBD, chosen because it is data-rich. The authors treat the outcome as an upper bound for this target class, not a general result across antigens.

Method

Three tasks map onto discovery-pipeline steps:

Sequencing was collected in scFv format; every submission was tested as IgG. Developability used a five-assay composite: pass at ≤3, fail at ≥4. SPR and single-point KinExA ranked clones similarly (Spearman ρ = 0.94).

Methods fell into four buckets: statistical heuristics, classical ML, attention or pretrained protein language models, and explicit 3D structure. Winners usually paired a PLM with a structure model.

Results

Challenge 1 is where AI looks real. About 13% of submissions gained ≥20-fold affinity over the parental antibody (to <10 nM). Aureka's AuraBind (structure-aware pairformer plus DPO) produced six developable antibodies below 10 nM, including a 94.7 pM clone statistically tied with the best experimental antibody at 113 pM. Six other groups made at least one developable <10 nM binder; eight groups made none. A non-AI consensus sequence reached 538 pM. Challenge 1 involved about 25 organizations and 119 submissions.

Challenge 2 is the sting. All but one AI ranking strategy underperformed random clone picking: 9.8%–13.8% of AI submissions beat the cluster control, versus 39% for random picks. Even cluster winners mostly landed at 11%–25%, still below random. No method won all three HCDR3 clusters. In cluster 27F, 58 submissions from 26 groups, 98.3% bound RBD and 24.1% reached sub-100 pM, but developability tracked the HCDR3 itself.

Challenge 3 is noisier. Across 23 organizations and 168 designs, 69.6% bound RBD and 30.4% did not; 14.3% reached sub-100 pM, 71.4% passed developability, and 53.6% were both binders and developable. Xencor's fine-tuned antibody LM with attention pooling produced the 2.9 pM winner. Nonbinders were more common than in experimental antibodies, and developability fell in the highest-affinity pool. Top designs sat 1–5 CDR substitutions from the given data, and every top submission reused an HCDR3 already present in the experimental set.

Why it matters

For people who design molecules, this is a cooler map than a launch blog: when affinity maturation is boxed in by sequencing, a leading model can skip 2–3 weeks of combinatorial library work. When the job is "find a better clone inside a cluster" or "invent CDRs outside the library," most models actively hurt. Random picks plus the most abundant clone beat most deep-learning pipelines on challenge 2.

Cross-task transfer is almost absent. Aureka won challenge 1; the same algorithm did not dominate challenges 2 and 3. There is no general-purpose antibody AI here.

Limitations

The authors are explicit: one antigen, organizers who also wrote the paper, blinding by trust rather than cryptographic isolation. Challenges 2 and 3 also handed out affinity labels that real early discovery does not have. The developability score is a sum, so one failed assay can be offset by others; they flag this as a scoring hole. Winning method details live mostly in the supplement. SARS-CoV-2 RBD is an overstudied target. Hit rates on a cold antigen would likely be lower.

Terms

Source

What people are saying

All paper explainers