Datology AI: Curated Pretraining Data Lifts 30B-A3B to 46.8% vs 37.7% Baseline
josh_wills · x · 2026-10-09
Datology AI reports that its curated pretraining data pushed a 30B-A3B model to 46.8% average benchmark score vs. 37.7% for the baseline, with gains persisting after post-training. A curated 12B-A1.4B model beat the 30B-A3B baseline despite using roughly one-fifth of the pretraining compute.
More from Research
- RNA fingerprinting method published in Cell infers drivers of cellular responses from Perturb-seq — anshulkundaje · 2026-10-10
- Team builds 8,000 sqft wet lab in 30 days; AI-generated compounds kill $13B pest Fall Armyworm — aadityabuilds · 2026-10-10
- Snorkel Scales Open Benchmarks Grants to $30M as Marin Lead Explains How Evals Guide Training — dlwh · 2026-10-10
- Computational biologists urge reconstructing key biological axes in latent space — anshulkundaje · 2026-10-10
- Success-Guided Sampling: sim-to-real RL nails dexterous assembly with zero demos — kevin_zakka · 2026-10-10
- AI-for-science debate: phenotype readouts matter more than gene-by-gene reconstruction — fabian_theis · 2026-10-10