Carbon-A: open model annotates 22k+ species genomes, yields 566M gene candidates
yb2698 · x · 2026-10-09
Hugging Face researchers released Carbon-A, an open model for discovering new genes in DNA. Using it, they annotated genomes from 22k+ species and produced 566 million gene candidates — about 16× the gene annotations in RefSeq — released alongside the model as a database.
Initial wet lab experiments support previously unannotated gene candidates, even in well-studied organisms. The team notes we can now sequence genomes faster than we can identify the genes in them; traditional gene finding relies on experimental data or comparison with well-studied species, leaving much of the tree of life poorly understood. They predict 2027 will be the year for biology, and want scientists to have capable open models without depending on the priorities of a handful of closed labs.
Related event: Open Model Carbon-A Discovers 566M Gene Candidates Across 22K Species(3 posts)→
More from Research
- Distilled influence embeddings make diffusion training data attribution a nearest-neighbor lookup — serrjoa · 2026-10-09
- DeepMind essay: what 15M Gemini interactions reveal about AI-accelerated science — JMateosGarcia · 2026-10-09
- ENPIRE turns coding agents into autonomous robotics researchers, hitting 99% success on dexterous tasks — chris_j_paxton · 2026-10-09
- ICLR 2027 asks ACs to catch reviewer conflicts they have no way to see — shaohua0116 · 2026-10-09
- Dietterich: treat OpenAI's math models like a deceased mathematician's notebook — tdietterich · 2026-10-09
- Paper: Naming a Fine Makes AI Agents Less Compliant, 46-Point Spread Across Models — niloofar_mire · 2026-10-09