110M DNA model matches 7B rival at finding prophages: domain-aligned pretraining data beats scale

bravo_abad · x · 2026-09-09

A comparison of genomic language models on prophage detection shows ProkBERT-mini-long—roughly 60x smaller than the 7B-parameter EVO2—achieves nearly identical genome-wide performance (0.700 vs 0.719), thanks to pretraining on prokaryotic DNA. The broader lesson for scientific foundation models: before scaling, check whether your pretraining data actually contain the domain you want the model to understand. Better domain alignment can be worth more than parameters.

Original post →

More from Research

Research channel →