Carbon-A: open model annotates 22k+ species genomes, yields 566M gene candidates

yb2698 · x · 2026-10-09

Hugging Face researchers released Carbon-A, an open model for discovering new genes in DNA. Using it, they annotated genomes from 22k+ species and produced 566 million gene candidates — about 16× the gene annotations in RefSeq — released alongside the model as a database.

Initial wet lab experiments support previously unannotated gene candidates, even in well-studied organisms. The team notes we can now sequence genomes faster than we can identify the genes in them; traditional gene finding relies on experimental data or comparison with well-studied species, leaving much of the tree of life poorly understood. They predict 2027 will be the year for biology, and want scientists to have capable open models without depending on the priorities of a handful of closed labs.

Related event: Open Model Carbon-A Discovers 566M Gene Candidates Across 22K Species(3 posts)→

Original post →

More from Research

Research channel →