Xaira says causal drug-discovery models need causal data, not just scale
Latent Space · rss · 2026-07-22
Latent Space discusses Xaira’s X-Cell drug-discovery model and why causal modeling needs causal data.
- The team found that scaling a model on a limited dataset eventually hit an information ceiling: test loss flattened while training loss kept improving.
- Their response was to build X-Atlas, a much richer dataset generated through large-scale CRISPR-based experiments.
- With roughly 30x more information, the model can continue scaling with both parameters and compute.
- The episode also covers why the team moved away from autoregression toward diffusion, how the system generalizes to real human-cell experiments, and why it beats previous linear baselines.
The central idea is that better biology models may depend less on parameter scaling and more on collecting causal, information-rich data.
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27