LabCompass finds swappable blood-cell recipes in four lab loops, lifting erythroid cells to 22%

2026-10-09

Guided flow matching designs CD34+ culture recipes. Across four loops and 78 million cells, one erythroid recipe reached 22%, with distinct recipes hitting the same fates.

What problem this solves

Pushing hematopoietic stem and progenitor cells toward one blood lineage is a dosing problem. Cytokines and small molecules interact, and the doses are continuous. Ten factors at four concentrations, crossed with two oxygen levels, is already millions of recipes. Exhaustive search is not a plan. Most virtual-cell models only run forward: given a recipe, predict the cells. Inverting that question, "which recipe produces this cell state," has mostly meant ranking a fixed catalogue of gene perturbations or discrete conditions. A catalogue cannot invent a dose it has never listed, and it scales badly once doses are continuous.

The same factor can expand cells or differentiate them depending on concentration and timing. Each recipe returns a mixture of cell states, not one cell type. The current gold-standard protocol for cord-blood megakaryocytes was distilled from 1,500 donor samples. For most hematopoietic fates it is still open whether one signal combination is enough, or several routes reach the same state.

Method

LabCompass treats recipe design as guided generation at inference time. A new target cell type does not require retraining. Designs are vectors of factor concentrations, oxygen, and days in culture. A conditional flow-matching model learns the prior over recipes that have actually been run: a velocity field pushes noise toward that distribution by integrating an ordinary differential equation.

At each step a one-step estimate maps the noisy state back to a clean recipe. A differentiable forward surrogate, itself a flow-matching model conditioned on the recipe, then generates a 26-dimensional cell state: 20 spectral-flow markers plus 6 scatter parameters. A readout turns those synthetic cells into a phenotype, usually a cell-type proportion or a marker mean and spread. The gradient of the loss against the requested phenotype is added to the generative velocity, so the trajectory is nudged toward recipes predicted to enrich the target while the prior keeps it near recipes that look experimentally plausible. Any forward model that is differentiable in the design variables can be plugged in.

Plain gradient descent on the recipe coordinates does not have that prior. It can leave the training support, follow degenerate directions of the surrogate, and collapse many initializations to one point. In simulation, LabCompass samples stayed inside the experimental bounds more often than gradient descent.

Selection is a fixed rule. Candidates are dropped if the predicted target proportion is too low, if oxygen falls outside 5-25%, if culture length falls outside 12-20 days, or if any axis exceeds the measured range (with a fractional margin). Survivors are ranked by predictive uncertainty, the standard deviation of the loss across repeated synthetic populations. For each lineage the lowest-uncertainty recipe is the exploit and the two highest are the explores. From loop 3, low-uncertainty candidates are also clustered, and representatives are taken from different clusters so the batch does not all sit on one mode. After each round of spectral flow, the forward model and the design prior are retrained from scratch. The forward architecture is re-picked by 25 random hyperparameter trials. The cell-type classifier is updated only when the annotation changes.

The wet lab used primary human cord-blood CD34+ cells in a basal medium plus about 10 cytokines and small molecules, read out on a 20-color spectral flow panel. An initial screen covered 116 designs and about 50 million cells. A loop is model proposal, selection, culture, measurement, and refit. Loop 1 was in silico only: six measured designs were held out, and no generated recipe was cultured. Loops 2 to 4 cultured model-proposed recipes, 6 to 16 designs per round, each round from a different donor. The abstract puts the spectral-flow dataset across those loops at about 78 million cells.

Results

On a synthetic grid at resolution 5, LabCompass beat uniform random search by a factor that grew from 1.73 in 2 dimensions to 97.91 in 30. LabCompass used a fixed budget of surrogate evaluations. Random search had to spend more evaluations to match it. The gap widens as the design space gets larger.

The hematopoietic loops then test whether that advantage survives real cells.

Loop 1 asked whether a known recipe could be retrieved. Design 210, taken from the literature, produced about 70% megakaryocyte progenitors (MgkPro). No training design exceeded 16%. The model had not seen design 210, yet generated concentrations near its signature doses, UM171 around 70 nM and butyzamide around 100 nM. Predictive uncertainty fell with Euclidean distance to design 210 (Spearman ρ = -0.72). After the six held-out designs were folded back in, the largest uncertainty drops landed where the model had been least sure (Spearman ρ = 0.702).

Loop 2 was the first culture of generated recipes, aimed at erythroid progenitors, phenotypic HSCs, and pre-pro-B cells as well as MgkPro. Erythroid conditions measured 6.1% and 7.4%, inside the top 16.3% of all tested designs. Megakaryocyte conditions measured about 44% and 10%, inside the top 3%. Phenotypic HSCs and pre-pro-B cells barely moved, which the authors attribute to missing medium components. After the update, uncertainty on held-out conditions fell 18.4% for MgkPro and 66.1% for erythroid progenitors. Pearson correlation of predicted and measured marker means rose from 0.437 to 0.873. Correlation of standard deviations was not significant before the update and reached 0.872 after it.

Loop 3 added two axes, M-CSF and a lymphoid cocktail built around IL-7 and DLL4 beads. Zeroing M-CSF in generated monocyte recipes and rescoring them dropped the predicted proportion almost in lockstep with the original prediction (Pearson r = 0.998 for CD14+ and 0.918 for CD16+). The high-uncertainty design that actually used the lymphoid cocktail, design 275, raised pre-pro-B and pro-B cells and CD19. The low-uncertainty design that did not use the new factor, design 274, did not. The model can exploit a cytokine only after someone puts it in the design space.

Loop 4 raised thrombopoietin above the previous ceiling of 2.5 ng/mL. The forward model predicted megakaryocyte-erythroid proportions would plateau beyond 25 ng/mL, and a titration agreed. Design 266, at 8.13 ng/mL TPO, brought the erythroid population to 22%, about 1.6 times control design 247, and ranked among the strongest CD235a conditions measured.

Clustering also separated recipes with similar predicted mixtures and different ingredients. One megakaryocyte-erythroid route looked like design 210: UM171 plus high butyzamide. The other used TPO plus IL-3, no UM171, and low butyzamide. Designs 261 and 263, the lowest-uncertainty members of those clusters, produced cumulative megakaryocyte and megakaryocyte-erythroid progenitor fractions whose means sat between 20% and 30%, in the same band as design 279, an independent biological repeat of design 210. Each point is one of five technical replicates. For CD16+ monocytes, two clusters differed mainly in LDL and UM171. Designs 267 and 268 measured 7.8% and 8.7%, above the positive control (design 277, a repeat of design 247) and above the model's own predicted band of 2-4%.

The same machinery was pointed at a published single-cell time course of human cytomegalovirus infection from Hein and Weissman. There the "recipe" is a viral expression state and the "cell state" is the host transcriptome. Guiding generation recovered held-out infection-stage phenotypes and, on the accompanying Perturb-seq screen, moved viral states in the direction of STAT2 and IFNAR2 knockdowns. The main text does not quote energy-distance numbers. Those comparisons are in the supplement.

Why it matters

For groups differentiating blood cells, growing organoids, or screening compounds, this is a closed loop that turns a requested cell-type fraction into the next culture conditions. Changing the target does not require a new training run, because the guidance happens at sampling time. The practical payoff is the alternative route: if a recombinant factor is expensive, unstable, or awkward for clinical manufacturing, a second recipe can aim at the same mixture.

The forward model does not have to be exact before it is useful. Masking M-CSF recovered a known monocyte dependency, and the TPO plateau looks more like a dose response than like linear extrapolation off the last measured point. The authors describe the surrogate as a world model of the culture and the guided generator as the policy that proposes the next experiment. The explore-versus-exploit rule used here is just a sort on uncertainty. A later loop could replace it with a proper active-learning schedule.

This is not an unattended lab. People still choose a handful of recipes to culture, and the forward model is retrained, with a fresh hyperparameter search, every round. Against Bayesian optimization and catalogue lookup, the specific addition is a generative prior over continuous doses that keeps proposals near recipes someone has already run.

Limitations

The paper states four limits. Performance depends on whether the training data contain the relevant signal, and a deep surrogate is a weak mechanistic explanation. Retraining the forward model every loop is expensive. Fine-tuning is mentioned as a future saving and was not used. A lineage whose required factor is absent from the design space cannot be reached, which is how the authors read the weak phenotypic-HSC result. Queries were one cell type at a time, and time-varying dosing schedules are left for later.

Other discounts are sharper. The loop 1 "recovery" of the 70% megakaryocyte recipe is an in-silico neighbor of a held-out condition. Nothing generated in that loop was cultured. When design 210 was repeated as design 279 in loop 4, it sat in the same 20-30% cumulative band as the two new routes. The paper does not reconcile that band with the original roughly 70% MgkPro measurement. Each round of 6 to 16 designs used a different donor, so recipe effects and donor effects are tied together. The authors note that instrument drift cannot be fully excluded, and they applied no batch correction.

Measured CD16+ monocytes at 7.8% and 8.7% beat the predicted 2-4%. The loop can still be worth running. A single predicted percentage is not a quantitative promise. The gradient-descent comparison is in silico only, on constraint satisfaction, not on cells produced. The cytomegalovirus study is a retrospective fit to published data, not a new experiment. Relative to a gold-standard protocol optimized across 1,500 cord-blood units, each round here is a small biological sample.

Terms

Source

What people are saying

All paper explainers